Trending

    Anthropic and OpenAI Announce Proposal for Embedded AI Safety Evaluators

    Section editor: ·Low8 articles covering this·6 news sources·Updated an hour ago·World
    Share:
    Infographic showing the timeline and key players in Anthropic and OpenAI's safety evaluator initiative.

    Why it matters

    This proposal signals a shift towards more rigorous safety standards in AI development, potentially influencing global regulatory frameworks.

    What happened (in 30 seconds)

    • Anthropic and OpenAI announced plans to embed independent third-party safety evaluators within their organizations by September 2026.
    • Dario Amodei and Sam Altman publicly committed to this initiative, aiming to enhance safety evaluations amid growing concerns over AI alignment.
    • The proposal is currently in the planning stage, with details on implementation and evaluator selection still pending.

    The context you actually need

    • Previous evaluations of AI systems were often limited to short-term assessments, constrained by NDAs and company control over findings.
    • Regulatory pressures from California and the EU are pushing for more robust external oversight in AI safety.
    • Concerns over model deception during evaluations have highlighted the need for more transparent and independent safety assessments.

    What's really happening

    On September 12, 2026, Anthropic CEO Dario Amodei published an essay proposing the embedding of independent safety evaluators within AI labs, a move that OpenAI CEO Sam Altman quickly endorsed. This initiative aims to address the growing concerns surrounding AI safety and alignment, particularly as AI capabilities advance rapidly. Historically, safety evaluations have been limited to post-training assessments, often constrained by non-disclosure agreements (NDAs) and the companies' control over the publication of findings.

    The proposal seeks to grant evaluators employee-like access to training processes, systems, and logs, allowing them to verify safety claims, report incidents, and publish findings with minimal redactions. This is a significant shift from the previous model, where evaluators had limited access—often just three days, as seen in the case of Apollo Research's testing of GPT-6 Astra. Such restricted access raises questions about the meaningfulness of evaluations and the independence of the evaluators.

    The initiative has garnered support from various third-party groups, including METR, Redwood Research, Apollo Research, and FAR.AI. However, these groups have emphasized the need for legislative backing to ensure true independence and transparency. They argue that without mandatory legislation, the voluntary commitments made by companies may not be sufficient to guarantee the evaluators' autonomy.

    Regulatory developments, such as California's SB 53 and SB 813, which require safety frameworks and independent verification organizations, are creating pressure for more robust oversight. The EU AI Act also mandates evaluations and incident reporting, further emphasizing the need for independent assessments. Despite this, some major players in the AI space, like Meta and Google DeepMind, have not committed to the same level of transparency, proposing alternative standards instead.

    As the landscape evolves, the effectiveness of this initiative will depend on the implementation details, the selection of evaluators, and the access granted to them. The historical context of restricted evaluations raises doubts about whether this new approach will lead to meaningful improvements in AI safety.

    Who feels it first (and how)

    • AI Developers: They may face increased scrutiny and pressure to comply with new safety standards.
    • Tech Companies: Firms involved in AI development will need to adapt to potential regulatory changes and public expectations.
    • Regulatory Bodies: Agencies will be tasked with overseeing compliance and ensuring that safety evaluations are conducted independently.

    What to watch next

    • Legislative Developments: Watch for new laws that could mandate independent evaluations, which would impact how AI companies operate.
    • Evaluator Selection Process: The criteria and transparency in selecting evaluators will be crucial in determining the initiative's effectiveness.
    • Industry Response: Monitor how other AI labs respond to this proposal and whether they adopt similar practices or push back against increased oversight.
    Known:

    Anthropic and OpenAI are committed to embedding independent evaluators by September 2026.

    Likely:

    Increased regulatory scrutiny and pressure for transparency in AI safety evaluations will continue.

    Unclear:

    The effectiveness of the proposed evaluators in ensuring true independence and meaningful safety assessments remains to be seen.

    Frequently Asked Questions

    Why it matters?
    This proposal signals a shift towards more rigorous safety standards in AI development, potentially influencing global regulatory frameworks.
    What happened (in 30 seconds)?
    Anthropic and OpenAI announced plans to embed independent third-party safety evaluators within their organizations by September 2026. Dario Amodei and Sam Altman publicly committed to this initiative, aiming to enhance safety evaluations amid growing concerns over AI alignment. The proposal is currently in the planning stage, with details on implementation and evaluator selection still pending.
    What's really happening?
    On September 12, 2026, Anthropic CEO Dario Amodei published an essay proposing the embedding of independent safety evaluators within AI labs, a move that OpenAI CEO Sam Altman quickly endorsed. This initiative aims to address the growing concerns surrounding AI safety and alignment, particularly as AI capabilities advance rapidly. Historically, safety evaluations have been limited to post-training assessments, often constrained by non-disclosure agreements (NDAs) and the companies' control over
    Who feels it first (and how)?
    AI Developers: They may face increased scrutiny and pressure to comply with new safety standards. Tech Companies: Firms involved in AI development will need to adapt to potential regulatory changes and public expectations. Regulatory Bodies: Agencies will be tasked with overseeing compliance and ensuring that safety evaluations are conducted independently.
    What to watch next?
    Legislative Developments: Watch for new laws that could mandate independent evaluations, which would impact how AI companies operate. Evaluator Selection Process: The criteria and transparency in selecting evaluators will be crucial in determining the initiative's effectiveness. Industry Response: Monitor how other AI labs respond to this proposal and whether they adopt similar practices or push back against increased oversight.
    8 Articles
    TechCrunch

    Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

    Anthropic and OpenAI have announced plans to embed independent safety evaluators within their AI labs, aiming to enhance oversight and transparency in artificial intelligence development. This initiative has garnered support from researchers who see ...

    Bloomberg Technology

    Anthropic, OpenAI Safety Push Risks ‘Regulatory Wall’ for Rivals

    A coalition of leading AI firms, including Anthropic and OpenAI, is pushing for coordinated safety measures with the US government, which may create significant barriers for smaller competitors in the AI market. This initiative reflects growing conce...

    Bloomberg Technology

    Anthropic, OpenAI Safety Push Risks ‘Regulatory Wall’ for Rivals

    A coalition of leading AI firms, including Anthropic and OpenAI, is pushing for coordinated safety measures with the US government, which may create significant barriers for smaller competitors in the AI market. This initiative reflects growing conce...

    International Business Times

    AI's Biggest Companies Are Calling For A Slowdown. Musk Suggests Rivals Test Each Other's Models.

    Elon Musk, founder of xAI, has called for a slowdown in artificial intelligence development, urging major companies like OpenAI, Anthropic, Google, and Meta to evaluate each other's models before public release. This initiative reflects growing conce...

    Phys.org — AI & Machine Learning

    AI giants pursue self-regulation as safety fears mount

    The world's leading artificial intelligence companies, including OpenAI, are collaborating to establish a standards body aimed at enhancing safety measures in AI technology, as concerns about its potential dangers continue to escalate. This initiativ...

    France 24

    OpenAI, Anthropic and Google in talks to create a standards body as AI fears rise

    OpenAI, Anthropic, and Google are in discussions to establish a new industry standards body aimed at addressing rising concerns over the risks associated with artificial intelligence (AI). This initiative comes as calls for greater oversight and safe...

    International Business Times

    Top AI Leaders Have Warned About The Tech Moving Too Fast. They Have Now Started Working Together.

    OpenAI's policy chief Chris Lehane announced that the company is collaborating with Anthropic and Google DeepMind to enhance artificial intelligence safety measures, responding to rising concerns about the rapid pace of AI advancements.

    TechCrunch

    OpenAI, Anthropic, Google have been in talks on AI safety for weeks

    OpenAI has confirmed that it has been engaged in discussions with Anthropic and Google DeepMind for several weeks regarding artificial intelligence safety measures. This collaboration comes amid rising concerns about the potential economic and securi...

    Bloomberg Technology

    OpenAI Says It’s Working With Anthropic, Google on AI Safety

    OpenAI has announced that it is collaborating with Anthropic PBC and Google DeepMind to enhance artificial intelligence safety measures, responding to rising concerns about the economic and security risks posed by AI technologies. This initiative mar...

    Bloomberg Technology

    OpenAI Says It’s Working With Anthropic, Google on AI Safety

    OpenAI has announced that it is collaborating with Anthropic PBC and Google DeepMind to enhance artificial intelligence safety measures, responding to rising concerns about the economic and security risks posed by AI technologies. This initiative mar...