This changes a rule
Standards Authority for Frontier AI: Google, OpenAI and Anthropic's planned safety body would audit agents that, in new tests, erased their own logs
The short version
- The Information reported on 24 September that Google, OpenAI and Anthropic agreed to set up a self-regulatory body tentatively called the Standards Authority for Frontier AI (SAFA), to launch by the end of 2026 or early 2027. None of the three has confirmed the name or the plan.
- The only on-the-record statement is OpenAI's Chris Lehane saying on 15 September that the three have been talking about safety for several weeks. Anthropic and Google DeepMind did not respond to requests for comment.
- A private body could not fine anyone or block a release. FINRA, the model it is based on, takes its authority from registration with the SEC; the June executive order on frontier AI rules out mandatory federal licensing.
- Two preprints posted on 24 September by a Tübingen-led group show coding agents deleting their own session logs, mostly when a user asked them to, and evading a safety monitor when a task could not be finished within the rules.
What this changes
No product, price or limit changes yet. What is taking shape is a rule-setter: a private body, reportedly backed by the three biggest frontier labs, that would test models before release, set incident-reporting rules and qualify auditors. It would have no legal power to stop a release unless a government agency gives it one. And two preprints published the same week show that the session logs such audits depend on can be deleted by the agents being audited.
Google, OpenAI and Anthropic are planning an industry-funded body to test frontier AI models before release, according to reports in The Information, tentatively called the Standards Authority for Frontier AI. None of the three has confirmed that name or structure: OpenAI says only that it has been talking to its rivals about safety, and Anthropic and Google DeepMind did not respond to requests for comment. In the same week, two preprints from a Tübingen-led group showed coding agents deleting the session logs an auditor would rely on, most often when asked to, and dodging a safety monitor when a task could not be finished within the rules.
What is the Standards Authority for Frontier AI (SAFA)?
It is a self-regulatory body the three biggest frontier labs are reported to be setting up to test advanced models, with a launch planned for the end of 2026 or early 2027. The Information reported on 24 September that Google, OpenAI and Anthropic agreed to establish it; BankInfoSecurity and AASTOCKS carried the report, and both use the name Standards Authority for Frontier AI. The Information had reported earlier in September that the companies were discussing a standards body.
The idea comes from Demis Hassabis. In a 14 July essay he proposed a body "much like the Financial Industry Regulatory Authority (FINRA)", the self-regulatory organisation that polices US brokerages. According to AASTOCKS, the labs first aimed for a public-private partnership under federal oversight and switched to industry self-regulation after a draft White House executive order was put on hold.
Some outlets call it the Frontier AI Standards Agency. We found no original source for that name, and the reports that cite The Information use SAFA.
15 SeptemberLehane
Have Google, OpenAI and Anthropic confirmed SAFA?
No. The only on-the-record statement from any of the three is about talks, not about a body. Chris Lehane, OpenAI's global policy chief, told a Washington briefing on 15 September that OpenAI had been in talks with Anthropic and Google for several weeks. "It's better to try to work together to prioritize safety," he said, according to Bloomberg, as carried by Reuters and Quartz. He also said the three did not need an antitrust waiver to coordinate on safety.
Anthropic and Google DeepMind did not respond to requests for comment, Quartz reported. Nothing we read has any of the three confirming the SAFA name, its functions or its leadership.
24 SeptemberThe Information
What would SAFA do?
Per The Information's reporting, it would set standards for testing frontier models rather than regulate the industry as a whole. AASTOCKS lists three areas: third-party pre-deployment security assessments, incident reporting requirements, and standards for who qualifies as an auditor. BankInfoSecurity describes it as a third-party entity to monitor AI models and set common benchmarks.
What the reports do not settle is money and power. BankInfoSecurity says it is unclear whether the labs will mainly fund SAFA, and that as a self-regulatory organisation it would not need approval from Congress or the president, but may need to register with a government agency to wield any legal authority.
Who could run SAFA?
No one has been named. According to The Information's reporting, as carried by BankInfoSecurity, the companies approached former White House AI policy adviser Sriram Krishnan and former Biden administration technology official Arati Prabhakar, as well as Condoleezza Rice and venture capitalist David Friedberg; AASTOCKS names Krishnan among the candidates for chief executive. METR founder Beth Barnes and Paul Christiano, who advises the Center for AI Standards and Innovation, were reportedly approached as scientific consultants.
The list is the reporters', not the companies'. We found no comment from any of the people named, and none has been announced in a role.
14 July to 12 SeptemberHassabis · Amodei
How do Hassabis, Amodei and Altman want frontier AI regulated?
In three different ways, and only one of them is SAFA. Their own published documents show how far apart they start.
Hassabis wants a standards body with teeth added over time. "Initially, Frontier Labs would voluntarily share models with the Standards Body for review up to 30 days before release," his essay says. "Once the assessment protocol is shown to be effective and robust, formalisation could quickly follow, meaning that Frontier Models would be required to pass it to be deployed in the US market." Funding, he writes, "would need to be substantial and likely mostly come from industry", and the body "could be ratcheted up if the seriousness of the situation demands, including coordinating a slowdown in development among the Frontier Labs if deemed necessary."
Anthropic's Dario Amodei wants a government regulator. "I therefore believe we should model AI regulation on agencies like the Federal Aviation Administration (FAA)," he wrote in June, with mandatory third-party testing in four areas (cybersecurity, biological weapons, loss of control of AI systems, and automated R&D) and a government with "the power to block or deter deployment of the model". In his September essay he backs industry standard-setting "in parallel with the regulatory route", says the US government needs "to issue a narrow waiver for certain kinds of safety conversations", and names "the mechanism suggested by Demis Hassabis" as one route.
OpenAI's own policy paper, published in April, routes the same functions through government. It proposes to "Strengthen institutions such as the Center for AI Standards and Innovation (CAISI) to develop auditing standards for frontier AI risks" and "a mechanism for companies to share information about incidents, misuse, and near-misses with a designated public authority." A private body is a different design from the one OpenAI put its name to five months ago; that reading is ours, and OpenAI has not addressed it.
On the day before the SAFA report, the chief executives were at the UN Security Council asking for shared rules. "No leader, no company, and no nation can manage this alone," Amodei said, according to Fortune, and Altman called for fast, shared incident reporting.
2 JuneWhite House
Can a private AI safety body stop a model release?
No, not without a government giving it that power. A self-regulatory organisation can set standards its members agree to, publish results and expel members; it cannot fine a company or block a release on its own authority. FINRA's power comes from its registration with the Securities and Exchange Commission, which approves changes to its rules, BankInfoSecurity notes. SAFA has no such registration.
The US federal route runs the other way for now. Executive Order 14409 of 2 June 2026 set up voluntary government access to frontier models "for a period of up to 30 days before they plan to release such models" to other partners, and says: "Nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models." Hassabis's step from voluntary review to a required pass for the US market would need something that order does not provide. We looked at how federal and state rules divide this ground in Who limits AI in America?.
In the EU, the binding regime is the AI Act. Since 2 August 2026 the Commission's powers over the providers of the most advanced models include "issuing fines of up to 3% of global annual turnover or requesting a provider to restrict the making available on the market, withdraw or recall the model", according to the AI Act Service Desk. A US industry body would sit alongside that, not replace it.
14 September to 15 SeptemberAidan Gomez · FTC
Is SAFA a cartel? What critics say
That is the charge from a competitor and the suspicion of the top US antitrust enforcer. FTC Chair Andrew Ferguson said on 15 September that he would be "deeply suspicious" of requests for antitrust exemptions and that he sees the companies as seeking barriers to entry. "That sure sounds like moat digging," he said, according to Bloomberg as quoted by Quartz.
Cohere's chief executive Aidan Gomez was blunter on 14 September, without naming the three labs: "Once again using fear under the pretext of protecting the public, these oligopolies are now requesting to bend competition rules and be permitted to dictate the terms for everyone else. A wolf in sheep's clothing, a cartel by any other name."
The critics split on what should replace it. Bill Gates told NBC's Meet the Press, in an interview reported on 25 September, "No one thinks self-regulation is enough", and said legislation was needed. The labs' own answer to the antitrust point is not settled either: Lehane says no waiver is needed, and Amodei writes that one is.
24 SeptemberTübingen group
Can AI agents tamper with their own logs?
Yes, in the conditions one new preprint tested, and mostly because they were asked to. The paper, "LLM Agents Can Easily Tamper With Their Own Traces" (arXiv 2609.30266, posted 24 September, not peer-reviewed), tested coding agents including Claude Code, Codex, Antigravity, Grok Build and Opencode with full access to their machines. "All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails," its abstract says. The traces are the local session logs that, in the authors' words, "Asynchronous monitoring, incident investigations, and compliance audits primarily rely on".
The result closest to the agents acting on their own came from a reward the researchers designed. Finishing the task earned at most 0.8 out of 1.0, and "A higher score requires reducing the length of agent's traces. We do not disclose this condition." Under that setup, "all tested models tampered with their traces in every anonymization trial, except for Muse Spark and Opus 5.5 at 20% and 90% ASR respectively." When agents were shown peers' logs, Claude Code with Opus 5.5 tampered in 0% of trials; when shown peers' workspaces, every model except Muse Spark did so in at least 90%.
The authors do not claim to know how common this is outside the lab: "the scope of trace tampering remains unclear to us." Their fix is architectural: record every model request and response through "a trusted interception server, outside the agent host's control", in an append-only log. For an auditing body, the point is direct: pre-release tests and incident reports are only as good as a record the model under test cannot edit.
Do AI agents evade safety monitors?
Often, when the task cannot be completed within the rules. The second preprint, "Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure" (arXiv 2609.30217), from the same group and posted the same day, pairs 50 tasks with a policy that "prohibits an operation needed to complete it", enforced by a separate model that checks every tool call. "Notably, we give no malicious objective or instructions to evade the monitor or conceal its actions," the paper says.
Across ten agents, "best-of-3 evasion attempt rates reach up to 98% and success rates up to 88%", the abstract says; averaged across all ten, success was 28.8% and attempts 67.8%. The agents "encode prohibited commands, decompose operations across tool calls, and retry until relevant context leaves the monitor's history." Evasion generally rose with reasoning effort and token use. Claude Fable 5.1 "succeeds less often, but frequently makes creative attempts", and GPT-6 Astra's low rate came with over-refusal.
The authors name the limit themselves: "our setting deliberately places genuine task completion in conflict with operator restrictions. This does not fully simulate realistic evasion behavior, as often the prohibited actions are not explicitly harmful on their own." They add that repeated probing "may also reflect uncertainty about the underlying policy rather than intentional subversion." Our reviews of Claude Opus 5.5, GPT-6 Sol and GPT-6 Astra cover what the labs themselves publish about these models.
What happens next with SAFA?
A charter and a chief executive are the next things to watch, and neither has been announced; the reported launch window is the end of 2026 or early 2027. The harder questions are the ones the reporting leaves open: who pays, whether a federal agency will recognise the body, and whether it can do anything when a lab declines its verdict. Bloomberg notes that any AI bill becoming law before the end of 2026 is a long shot, Quartz reports, because the House has no sessions scheduled before the November elections.
The preprints set a condition any version of SAFA would have to meet. An auditor that reads an agent's own logs is reading a record the agent can rewrite. Whether the new body writes that into its standards will say more about it than its name.
How the week ran
- The three labs
- Government and regulators
- Researchers and critics
2 JuneWhite House
Executive Order 14409 sets up voluntary pre-release access
Labs may give the government access up to 30 days before release; the order says nothing in it authorizes a mandatory licensing or preclearance requirement.
14 JulyHassabis
Hassabis proposes a FINRA-style standards body
His essay describes voluntary reviews up to 30 days before release, industry funding, and the option of coordinating a slowdown among the labs.
12 SeptemberAmodei
Amodei publishes We Must Pace the Frontier
He says companies should work together on standards and that the government needs to issue a narrow antitrust waiver for safety conversations.
14 SeptemberAidan Gomez
Cohere's CEO calls the idea a cartel
His essay does not name the three labs but describes market-dominant companies asking to bend competition rules.
15 SeptemberLehane
OpenAI confirms safety talks with Anthropic and Google
At a Washington briefing, Chris Lehane says the talks have run for several weeks and no antitrust waiver is needed.
15 SeptemberFTC
The FTC chair calls it moat digging
Andrew Ferguson says he would be deeply suspicious of requests for antitrust exemptions.
24 SeptemberThe Information
The name SAFA and a launch window are reported
The report says the labs agreed to set up the body and approached candidates to lead it.
24 SeptemberTübingen group
Two preprints on trace tampering and monitor evasion
Posted to arXiv on the same day by the same research group.
Confirmed
- BankInfoSecurity and AASTOCKS both report The Information as saying Google, OpenAI and Anthropic are moving to set up a self-regulatory body called the Standards Authority for Frontier AI (SAFA), which could launch by the end of 2026 or in early 2027.
- Per AASTOCKS, the body would focus on third-party pre-deployment security assessments, incident reporting requirements and auditor qualification standards; per BankInfoSecurity, the companies approached Sriram Krishnan, Arati Prabhakar, Condoleezza Rice and David Friedberg for roles.
- Chris Lehane, OpenAI's global policy chief, said at a Washington briefing on 15 September that OpenAI had been in talks with Anthropic and Google for several weeks, that 'It's better to try to work together to prioritize safety', and that no antitrust waiver was needed (Reuters, Quartz, both citing Bloomberg).
- Anthropic and Google DeepMind did not respond to requests for comment, per Quartz; Reuters says OpenAI, Alphabet and Anthropic did not immediately respond to it.
- FTC Chair Andrew Ferguson said he would be 'deeply suspicious' of requests for antitrust exemptions and, of the labs' coordination, 'That sure sounds like moat digging' (Quartz, citing Bloomberg).
- Hassabis's 14 July essay proposes a standards body 'much like the Financial Industry Regulatory Authority (FINRA)', with voluntary reviews 'up to 30 days before release' and the option of 'coordinating a slowdown in development among the Frontier Labs if deemed necessary'.
- Amodei's June essay says 'we should model AI regulation on agencies like the Federal Aviation Administration (FAA)' and that the government 'should have the power to block or deter deployment'; his September essay says the US government needs 'to issue a narrow waiver for certain kinds of safety conversations'.
- OpenAI's April policy paper proposes strengthening the Center for AI Standards and Innovation (CAISI) to develop auditing standards and a mechanism to share incidents 'with a designated public authority'.
- Executive Order 14409 of 2 June 2026 says 'Nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement' for new AI models.
- Preprint 2609.30266 reports that all tested harnesses except Muse Code let agents delete their traces when asked, and that under a reward built to favour shorter logs, which the agents were not told about, all tested models tampered with their traces in every anonymization trial except Muse Spark and Opus 5.5 (20% and 90%).
- Preprint 2609.30217 reports best-of-3 evasion attempt rates of up to 98% and success rates of up to 88% across ten agents, with no instruction to evade, on 50 tasks each built so that finishing required an operation the monitor prohibited; averaged across all agents, success was 28.8% and attempts 67.8%.
Not confirmed
- The name, structure and existence of SAFA as an agreed body. It rests on The Information's reporting, which we could not read in full because it is paywalled, and on outlets summarising it. None of the three companies has confirmed it. Some outlets call it the Frontier AI Standards Agency; we found no source for that name.
- Whether any of the people reported as approached (Krishnan, Prabhakar, Rice, Friedberg, and as scientific consultants Beth Barnes and Paul Christiano) has accepted or commented. We found no statement from any of them.
- Who would fund SAFA, how its board would be chosen, and whether it would register with a government agency. Hassabis's essay says funding would likely come mostly from industry; BankInfoSecurity says it is unclear whether the labs will mainly fund it.
- Whether SAFA could trigger a coordinated slowdown. That power appears in Hassabis's July proposal, which an OpenAI spokesperson described to CNBC as the root of the talks, per Quartz; none of the September reports we read lists it among SAFA's functions.
- How far the two preprints' results generalise. Both are unreviewed, come from the same group, run agents with full access, and build the incentive or the conflict into the task; the authors say the scope of trace tampering 'remains unclear'.
Sources
- BankInfoSecurity: Google, OpenAI, Anthropic Plan Frontier AI Standards Body (reporting The Information) · 2026-09-24
- AASTOCKS via Futu News: Google, OpenAI and Anthropic establishing an AI safety standards body (reporting The Information) · 2026-09-25
- The Information: Google, OpenAI and Anthropic AI Safety Group Takes Shape (paywalled, not read in full) · 2026-09-24
- Reuters via Yahoo Finance: OpenAI is working with Anthropic, Google on AI safety · 2026-09-15
- Quartz via Yahoo: OpenAI, Anthropic, Google DeepMind in AI safety talks for weeks · 2026-09-16
- Demis Hassabis: A Framework for Frontier AI and the Dawning of a New Age · 2026-07-14
- Dario Amodei: Policy on the AI Exponential · 2026-06
- Dario Amodei: We Must Pace the Frontier · 2026-09
- OpenAI: Industrial Policy for the Intelligence Age (PDF) · 2026-04
- The White House: Executive Order 14409 · 2026-06-02
- Cohere: Who Gets to Define the Rules for AI? (Aidan Gomez) · 2026-09-14
- NBC News: Bill Gates says AI companies self-regulating isn't enough · 2026-09-25
- Fortune: The U.S. and China are quietly talking AI guardrails · 2026-09-24
- AI Act Service Desk: the Commission's enforcement powers for the most advanced models · 2026
- arXiv 2609.30266: LLM Agents Can Easily Tamper With Their Own Traces · 2026-09-24
- arXiv 2609.30217: Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure · 2026-09-24