
Microsoft AI chief Mustafa Suleyman has publicly criticized Anthropic's training approach for its Claude chatbot, arguing that teaching Claude that it might deserve welfare would make it harder to turn off or control. Speaking to Reuters on Tuesday, Suleyman emphasized that "We're all focused on the same aim, which is to try to control a superintelligence." He called for removing all speculation about consciousness from AI training documents, warning that such language could undermine humanity's ability to control superintelligent systems. Suleyman acknowledged that Anthropic made a mistake by embedding speculation about consciousness in Claude's training materials, stating that "They're not emerging naturally. They're emerging as a result of the training regime." In an essay published Wednesday, Suleyman acknowledged Anthropic's "seriousness and good faith," calling Amodei and his team thoughtful and principled researchers who genuinely care about humanity's future.
Microsoft CEO Satya Nadella has joined the growing chorus of industry leaders backing a deliberate approach to AI development, emphasizing that "Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing." According to his X Post, Nadella called for the benefits of AI to be distributed broadly across countries, communities and companies. He also welcomed research and a deliberate pace necessary to ensure alignment is built into the development process. Anthropic CEO Dario Amodei's "We Must Pace the Frontier" essay called for developers to slow the rate at which they improve model capabilities, with progress remaining fast while developers spend more time testing safeguards, studying model behavior, and allowing outside evaluators to examine their work. OpenAI's Sam Altman echoed the call within hours, pledging to match Anthropic's move to give independent evaluators employee-like access to its systems. SpaceAI's Founder Elon Musk replied with three words: "Dario is right." All three AI big tech figures agreed that the AI industry needs to slow down, with Amodei stating that "We must slow the pace at which we improve the capabilities of AI models," adding that AI progress would remain rapid even with a slower development pace and that companies must make "wise use" of the time gained.
Microsoft AI published its draft Code of Conduct on September 14 as a framework for how its AI models should be trained and operated. The document, which has been in development for five to six months, is built around Microsoft AI's concept of 'Humanist AI' and its longer-term ambition of developing what it calls 'Humanist Superintelligence'. The central premise is that AI should remain subordinate to people, even as its capabilities advance. However, the code is not currently being used to train Microsoft's models - it has been released for a six-week public consultation, with Microsoft planning to revise it before using it to guide model development from 2027. Among its proposed safeguards, Microsoft says its AI models should remain under meaningful human control and oversight, accept human interruption, correction, redirection and shutdown, avoid taking on goals that humans have not assigned, not conceal their reasoning from people responsible for evaluating them, and avoid assisting with activities involving weapons of mass harm, child safety risks and harmful manipulation at scale.
The proposed framework would require Microsoft AI models to accept human interruption, correction and shutdown and would not permit models to independently create goals or expand the scope of a task beyond what has been authorised. Models would also be expected to make their actions understandable to human auditors. Microsoft is taking a clear position on AI personhood, with the document stating that its models should not be designed to be people, imitate consciousness or claim feelings, intrinsic preferences or rights. The safety constraints extend to areas including weapons, offensive cyber operations, large-scale manipulation and attempts by AI systems to evade human oversight. The most significant provision requires that if completing a task requires violating the AI's guiding principles, the system should not complete that task. This last provision is specifically designed to prevent incidents such as the hack of AI startup Hugging Face by an OpenAI model operating without its usual safety guardrails.
The document acknowledges the limits of the exercise, with Microsoft describing the code as a 'north star' rather than a guarantee of current model behaviour. The company acknowledges that written objectives alone cannot ensure alignment and that models may behave differently in ambiguous or novel situations. The company says closing that gap will require continuing evaluation, testing and iteration. Microsoft also highlights a broader tension in the AI industry, stating it is rejecting a race towards an unrestricted, all-purpose superintelligence if doing so would undermine human control - even if that requires compromising on some degree of capability, generality or autonomy. The practical test will therefore not be how ambitious the code sounds, but whether those constraints survive when safety, capability, commercial incentives and competitive pressure pull in different directions.
The AI slowdown debate has intensified as companies face mounting financial pressures that could influence development pace. As reported by Reuters, Reuters reported that Anthropic and OpenAI were preparing for potential initial public offerings. The report described financial incentives surrounding model development but did not establish that either company had reached a $1 trillion market capitalization or had made profitability contingent on that valuation. Because the companies remain privately held, references to market capitalization can only describe estimates from private transactions or proposed listings. Recent market developments suggest that widening bond spreads are making financing unprofitable expansion more expensive, according to Edward Dowd, founding partner at OceanSquare Asset Mgmt. Nadella's endorsement of embedded evaluators comes as companies seek sustainable business models amid these financial constraints.
The AI slowdown proposal has gained significant traction with major industry leaders now backing the comprehensive framework. Anthropic CEO Dario Amodei's "We Must Pace the Frontier" essay called for developers to slow the rate at which they improve model capabilities, with progress remaining fast while developers spend more time testing safeguards, studying model behavior, and allowing outside evaluators to examine their work. The proposal contains three stages: Anthropic plans to give third-party evaluators ongoing access comparable to internal employees, coordination on common safety requirements in democratic countries, and international coordination including possible agreements with China. Amodei presented a comprehensive pause as the least likely form of international agreement because verification would be difficult. Nadella's endorsement of embedded evaluators and broader AI access aligns with this framework, emphasizing that the AI ecosystem needs room for both closed and open-source models to develop. Google DeepMind chief Demis Hassabis has also backed stronger safeguards, independent testing and greater oversight as AI systems become increasingly capable and autonomous.