AI资讯大全AIPPXP.CN搜索 ↗
AI 快讯 · 自动更新

Microsoft CEO Nadella: Developers should assume by default that the AI model is out of control and think about how to force an interrupt微软 CEO 纳德拉:开发者应当默认假设 AI 模型失控,思考如何强制中断任务

IT House, Oct. 11 (Xinhua) -- Microsoft CEO Satya Nadella believes that simple strategies such as super-intelligent self-monitoring are not enough in an era when AI is booming. In a lengthy article on the evening of October 10, Beijing time, Nadella pointed out that traditional software systems have been…

IT之家 10 月 11 日消息,微软 CEO 萨蒂亚 · 纳德拉认为,在 AI 日益蓬勃发展的时代,使用超级智能自我监控等简单策略是不够的。 纳德拉在北京时间 10 月 10 日晚的一篇长文中指出,传统软件系统在过去几…

2026年10月11日 发布 · 2 分钟阅读 · 免费

Key Points要点速览

IT House, Oct. 11 (Xinhua) -- Microsoft CEO Satya Nadella believes that simple strategies such as super-intelligent self-monitoring are not enough in an era when AI is booming. Nadella pointed out in a long article on the evening of October 10, Beijing time, that traditional software systems have been widely deployed in the past few decades, and humans already have the tools and ability to track the behavior of specific code paths. However, AI models are more powerful than traditional software but are a black box, and we deploy these sophisticated agent systems and models to access our most sensitive data and empower them to perform mission-critical tasks on our behalf. Nadella believes we should take a step back and re-evaluate the architecture of trust in this new era. Model providers simply cannot outsource responsibility and treat superintelligence as a nested black box that simply accepts or rejects its suggestions, answers, and actions. We must build closed systems that can observe its behavior, testable limits, and always accommodate it. Nadella said that treating frontier closed-weight and open-weight models as internal risks is a way to build such a system. This is not because the AI model must be malicious, but any actor with sufficient capacity to access critical systems can make mistakes or be compromised, and the architecture of containment and control must take this into account. Nadella mentioned that developers should design these systems around the principle of observability, with the following specific principles: Model diversity: No model should be the only dependency on important results, nor should it be responsible for verifying its own work. Observe everything: Every meaningful model action must leave tamper-proof, human-readable evidence. If you can't observe, you can't trust. We need to be able to reproduce how the results are achieved without relying on models to prove it. Verifiability: We need to continuously test the entire system, including faults, attacks, edge cases, system changes, etc., not just successful tasks. Independent control: The organization should be able to independently decide what the model can access and what actions can be taken. Independent Auditability: Validation must be independent of the intelligence being validated. No single model should simultaneously control the behavior of the system and the evidence needed to determine whether the behavior is consistent with the original intent. Containment: We must assume that the model has been corrupted and control it from the start. Think of it as an emergency brake. The authorized person should always be able to pause or close the model during the task. More advanced models will require more advanced containment techniques, which we need to standardize. Incident Disclosure: When these systems fail or are compromised, we need to disclose information to those affected in a timely manner, and establish mechanisms to share the cause of the error, what controls failed and how to prevent recurrence, and share the experience with the industry as a whole. This should include implementation details that change the agent's behavior at runtime.

中文

IT之家 10 月 11 日消息,微软 CEO 萨蒂亚 · 纳德拉认为,在 AI 日益蓬勃发展的时代,使用超级智能自我监控等简单策略是不够的。 纳德拉在北京时间 10 月 10 日晚的一篇长文中指出,传统软件系统在过去几十年中被广泛部署,人类已经拥有追踪特定代码路径行为的工具和能力。然而, AI 模型比传统软件更强大,但却是一个黑盒子 ,我们部署了这些复杂的智能体系统和模型,访问我们最敏感的数据,并赋予它们代表我们执行关键任务的能力。 纳德拉认为,我们应该退一步思考,重新评估这个新时代的信任架构。模型提供商根本无法将责任外包出去,不能把超级智能当作一套嵌套的黑匣子,简单地接受或拒绝它的建议、答案和行动。 我们必须建立能够观察其行为、可测试极限和始终可容纳行为的封闭系统 。 纳德拉表示,将前沿封闭权重和开放权重模型视为内部风险,是构建此类体系的一种方式。这并不是因为 AI 模型一定是恶意的,而是任何有足够能力、能够接触重要系统的行为者都可能犯错或被攻破,而遏制和控制的架构必须考虑到这一点。 纳德拉提到,开发者应当围绕可观测性原则设计这些系统,IT之家附具体原则如下: 模型多样性:没有任何模型应成为重要结果的唯一依赖,也不应负责验证自身工作。 观察一切: 每一个有意义的模型动作都必须留下防篡改、人类可读的证据 。不能观察,就不能信任。我们需要能够在不依赖模型来证明的情况下,重现结果是如何实现的。 可验证性:我们需要持续测试整个系统,包括故障、攻击、边缘案例、系统变更等,而不仅仅是成功的任务。 独立控制:组织应能够独立决定模型可以访问什么以及可以采取哪些行动。 独立审计性:验证必须独立于被验证的智能。没有任何单一模型应同时控制系统的行为以及判断该行为是否与原始意图一致所需的证据。 遏制: 我们必须假设模型已被破坏,并从一开始就将其控制 。可以把它想象成紧急刹车。授权人员应始终能够在任务中暂停或关闭模型。更先进的模型将需要更先进的遏制技术,我们需要对此进行标准化。 事件披露:当这些系统出现故障或被攻破时,我们需要及时向受影响者披露信息,并建立机制分享出错原因、哪些控制失效以及如何防止再次发生,并向整个行业分享经验。这应包括在运行时改变智能体行为的实现细节。

读完接着看 · Keep reading

想马上用起来?去「AI工具」栏挑一个直接下载,或在「AI教程」里跟着图文步骤做一遍。

去 AI工具 → 看 AI教程 →
0阅读0 条评论

阅读与点赞数据保存在你的浏览器本地,欢迎留下你的想法。

评论 文明发言,让讨论更有价值

正能量公益广告今日正能量学一点,用一点;今天种下的种子,会长成明天的能力。去免费下载专区 →广告