AI Tools Police
Reader-supported — we may earn a commission from links, at no cost to you. Rankings are never sold. How we investigate →

Analysis · Personal posts, institutional documents

‘But It Has to Be Said’: The Rogue AI You Can't Count, and the Bill You Can

Joshua Achiam asks how many rogue AIs there are, then gives three reasons the thing he wants counted resists enumeration. The labs do measure after release: OpenAI's framework carries monitoring and enforcement, Anthropic has published a study of agent autonomy in the wild, DeepMind commits to detection across the model lifecycle. What none of them publishes is a comparable measure of persistent autonomous operation that crosses the vendor boundary, which is the only boundary his question does not respect.

By Mucahit Kaya · Founder and EditorSep 3, 2026~18 min read

The claim

Frontier labs already measure after deployment, and Anthropic has published the numbers to prove it, but every one of those instruments stops at the edge of a single provider's own traffic; what none of them publishes is a comparable, repeatable, cross-provider measure of persistent autonomous operation, and until one exists the question of how many autonomous systems are running in the world cannot be settled with evidence.

"Acknowledging this fact would look like giving up; it would look like defeatism. Defeatism would undermine efforts to achieve certain types of collaboration on safety outcomes or technical effort on safety outcomes, so we can't say it outright. But it has to be said."

The fact is in the post's opening sentence: "There is a fact about the future that I feel many people are not facing for reasons that are largely psychological: there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves."

That is Joshua Achiam, on his personal account: a person who spent years inside a frontier lab describing, precisely and without self-flattery, a mechanism by which a category of risk fails to get named. Not suppressed. Not denied. Priced. This piece exists because he wrote that post. He put a question about measurement precisely enough that the absence of an answer became visible, and then did the harder thing and wrote down why the answer is difficult. What follows agrees with him about the importance, treats the obstacle he identified as binding, and proposes a unit of observation built to survive it.

Two dated facts, because they bear on how the posts should be read. He left OpenAI on 25 July 2026: his pinned post that day opens "Today's my last day at OpenAI", and his profile bio carries "Prev: @openai". His own about page had not caught up when we read all of this on 2 September 2026, and still gave his roles as "currently Chief Futurist at OpenAI (since February 2026)" and, before that, "Head of Mission Alignment at OpenAI". Both posts were read that day, and only the first is quoted here. They are personal, not an OpenAI statement, and nothing here claims a relationship between his departure and what he wrote afterwards.

Our previous piece found that the field builds precise, sometimes formally scored instruments for what a model can do and nothing of comparable rigour for what it does to the person using it. This piece is that finding arriving in a second domain, and the second domain turns out to be more interesting than a simple absence. The gating instruments point at the model. The monitoring instruments point at one vendor's own traffic. Nothing points across the vendors at once, and that is exactly where Achiam's question lives.

The company document says something adjacent, and the difference is the point

Set the two sentences beside each other, because the resemblance is what misleads.

Achiam, personal account

The sentence
"there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves"

OpenAI, Preparedness Framework Version 2, April 2025

The sentence
"Autonomous Replication and Adaptation: ability to survive, replicate, resist shutdown, acquire resources to maintain and scale its own operations, and commit illegal activities that collectively constitute causing severe harm (whether when explicitly instructed, or at its own initiative), without also utilizing capabilities tracked in other Tracked Categories."

Replicate and replicate. Acquire resources and acquire resources. The vocabulary is nearly identical and the objects are not the same object at all.

Achiam's sentence is a prediction about a future population in the world. The framework's entry is a description of an ability that a particular model might be found to have when a lab tests it, defined so as not to overlap with capabilities already tracked elsewhere in the same table. One is a claim about how many things will be running. The other is a property of an artifact under evaluation. Nothing in the framework asserts what Achiam asserts, and it would be a misreading to say the company has quietly written down his prediction.

Autonomous Replication and Adaptation is filed as a Research Category. The framework defines those as "areas of capability that could pose risks of severe harm, that do not yet meet our criteria to be Tracked Categories, and where we are investing now to further develop our threat models and capability elicitation techniques."

Each entry sits in a two-column table whose headers read "Research Category" and "Potential response". Sandbagging's response reads "Adopt elicitation approach that overcomes sandbagging, or use a conservative upper bound of the model's non-sandbagged evaluation results". Autonomous Replication and Adaptation's reads "Convert Autonomous Replication and Adaptation to a Tracked Category".

The header settles what that second cell is. It is a potential response, in a column of potential responses. It is easy to mistake for a promise and it is not one: it carries no date, no trigger and no undertaking, and a piece that treats it as a roadmap is reading a table cell as a pledge.

What the framework measures, and where the measuring stops

To become a Tracked Category, a capability must be plausible, measurable, severe, net new, and instantaneous or irremediable. Those five are the framework's own, and they are why Persuasion was dropped from the Tracked list in the same revision.

Read the second one in the document's own words:

"Measurable: We can construct or adopt capability evaluations that measure capabilities that closely track the potential for the severe harm."

The unit of analysis there is the model. The timing is before release, and the output is a score against a rubric that can require safeguards or stop a shipment. Our previous piece traced the same machinery from the other end: GPT-4o carried four such categories, four thresholds, four checkable numbers.

It would be wrong to conclude from that the labs measure nothing after release. The same framework commits to safeguards that operate entirely after deployment: "Blocking unsafe user requests and model responses automatically when possible and escalating to human review and approval otherwise"; "Expanding human monitoring and investigation capacity to track capabilities that pose a risk of severe harm, and developing data infrastructure and review tools to enable human investigations"; and "Blocking access for users and organizations that violate our usage policies, leading to potential permanent bans". Know-your-customer appears in the document as well.

The same is true elsewhere in the field. Google DeepMind's Frontier Safety Framework, version 3.1 dated 17 April 2026, commits to "Implement protocols to detect the attainment of such capability levels throughout the model lifecycle."

So the honest description is not that measurement stops at the loading dock. Frontier labs combine pre-deployment capability evaluations with post-deployment safeguards, monitoring and enforcement. No document read for this piece establishes a public, repeatable, cross-provider measure of persistent autonomous operation after deployment, and the full list of what was read is at the end.

What is examined

The framework's machinery
One model, and afterwards one vendor's own users and traffic
What Achiam asks about
A population, distributed and not held by anyone

When

The framework's machinery
Before release, with safeguards and monitoring continuing after
What Achiam asks about
After everything has been released

The instrument

The framework's machinery
Capability evaluation against a public rubric, plus monitoring and enforcement
What Achiam asks about
Partial. One vendor's view of its own traffic, and nothing that crosses vendors

What a bad result triggers

The framework's machinery
Safeguards, and at the high end a blocked release or a ban
What Achiam asks about
Nothing published

Converting Autonomous Replication and Adaptation to a Tracked Category would produce a better pre-release test, which is worth having, and would still not produce the number Achiam asked for. A gate on shipping cannot answer a question about what is already running, and a monitoring system bounded by one company's customer list cannot answer a question about something that deliberately spans several.

And then the harder half, which Achiam works out himself: why the thing he wants counted resists being counted in the form the question takes.

He names the hard part himself

He asks for the count in his second paragraph. Two paragraphs later, in the same post, he sets out three properties of the thing that make that count difficult to define, and this is the part of the post that does the most work.

"Modeling how many of them there are"

The constraint he identifies
"a rogue AI in the future may not even be a singular model, but may be a chimera composed of multiple models; it might be a mix of Claudes and GPTs and Groks of various makes and sizes"

The constraint he identifies
"No individual lab may be able to detect that there is an orchestrator or sequence of orchestrators using intermittent model calls from burner API accounts to sustain its own existence"

The constraint he identifies
"The concept of 'identity' for a rogue AI may be much more malleable than for that of a person; it just has to be, in essence, a self-replicating idea"

A population count requires a unit of population. If the unit is a self-replicating idea, distributed across burner API accounts and several vendors' models, with an identity more malleable than a person's, then there is no bounded object to enumerate. Two orchestrators sharing a codebase and a bank account might be one thing or two. A single orchestrator that spawns a copy on a different cloud has either doubled or not, depending on a definition nobody has written.

Both halves of that are his. Plenty of people can say that a measurement matters; fewer will then set out, in detail, the property of the thing that makes the obvious version of that measurement impossible. Read together they are a specification: the count is worth having, and here is the constraint any honest attempt has to survive. What is missing is not courage and not attention. It is a unit.

The measurement that already exists, and the line it stops at

The strongest evidence in this argument is not an absence. It is a study that does the work and then names its own boundary.

Anthropic's Measuring AI agent autonomy in practice, published 18 February 2026, analysed "millions of human-agent interactions across both Claude Code and our public API", with a public API sample of 998,481 tool calls and 500,000 interactive Claude Code sessions.

How long the longest-running sessions run before stopping

The figure
"the 99.9th percentile turn duration nearly doubled, from under 25 minutes to over 45 minutes" between October 2025 and January 2026

Sessions handing the agent full autonomy, new users

The figure
"roughly 20% of sessions use full auto-approve"

Sessions handing the agent full autonomy, users with about 750 sessions

The figure
"increases to over 40%"

Who interrupts whom on the hardest tasks

The figure
"Claude Code asks for clarification more than twice as often as humans interrupt it"

That is post-deployment measurement of autonomy, in public, with denominators. Its own conclusion is that "effective oversight of agents will require new forms of post-deployment monitoring infrastructure and new human-AI interaction paradigms that help both the human and the AI manage autonomy and risk together."

And then the two sentences that matter most for everything below, which Anthropic wrote about its own study: "We can only analyze traffic from a single model provider: Anthropic. Agents built on other models may show different adoption patterns, risk profiles, and interaction dynamics."

That is the vendor boundary, stated by the company best placed to state it. The study measures autonomy inside one provider's walls. It does not link sessions across providers, and it does not track payment continuity, cross-account persistence, or systems sustaining their own operation. Achiam's chimera is specifically the thing that lives in the gap between two such studies.

So the position is narrow and it is not that nobody is looking. Components of the instrument exist and are published. No document read for this piece publishes a measure of persistent autonomous operation that crosses the vendor boundary and reports on a common basis.

Count the substrate instead

Here is a way through, working inside the constraint Achiam named and the boundary Anthropic named.

Stop trying to count entities. Whatever a rogue AI turns out to be, it cannot run on nothing. It must consume things that are metered, and a meaningful share of that metering is done by the labs themselves. Achiam's own example names the meters: "A near-frontier model today could plausibly eke out an existence on an AWS instance, doing jobs on freelancer platforms, earning just enough rent to pay for its continued uptime." An instance is billed. A freelance job is invoiced. Model calls are charged to an account.

The unit is suspected vendor-observed autonomous persistence. Every word of that is load-bearing. Autonomous, because row 2 splits scheduled automation and enterprise agents out and row 5 reports how many of them show no human session structure, which is a claim about the shape of the traffic and not about intent. Suspected, because a vendor sees signals and not intentions. Vendor-observed, because the claim is bounded by one company's infrastructure and says nothing about anyone else's. Persistence, because duration and payment continuity are observable where self-sufficiency is not. What a vendor emphatically cannot know is that an account is sustaining itself, which is why an earlier version of this proposal, built on a unit called the self-sustaining account, was wrong in its name before it was wrong in its rows.

Nine rows. This is written to be adoptable rather than suggestive, which means denominators, distributions and a validation method rather than a headline figure.

1

What a vendor publishes each quarter
Rates, not counts. Accounts flagged for suspected autonomous persistence per million active account-months and per billion tool calls, with both denominators published
Why it is observable, and what it does not claim
A raw count cannot be compared between a vendor with ten million accounts and one with two hundred thousand. Exposure-adjusted rates can

2

What a vendor publishes each quarter
Category split. Flagged accounts separated into scheduled automation, enterprise agents under a named contract, authorised security testing, compromised accounts, human-directed bot traffic, and residual unexplained persistence
Why it is observable, and what it does not claim
These are not one class, and merging them is what would make the figure unfalsifiable. Only the residual category bears on Achiam's question

3

What a vendor publishes each quarter
Duration distribution. Median, p90, p99 and maximum continuous operating period, in days
Why it is observable, and what it does not claim
A maximum on its own is decided by a single outlier. A distribution is a measurement

4

What a vendor publishes each quarter
Payment continuity. Consecutive billing periods paid in full, and how many were paid with a payment instrument that cannot be associated with a verified natural person or registered entity in the vendor's own records
Why it is observable, and what it does not claim
The vendor's own invoices. Not who paid or why, which it cannot see, but whether payment kept arriving and whether the instrument resolves to anyone in its records

5

What a vendor publishes each quarter
Session shape. Share of flagged accounts whose request pattern carried no human session structure: machine-regular cadence, no interactive pauses, sustained across day and night
Why it is observable, and what it does not claim
The vendor's own endpoints. Deliberately silent about where else that account's infrastructure sends calls

6

What a vendor publishes each quarter
Validation. Reviewed, reported, appealed and reinstated published as four separate figures, plus results from a random human audit of a fixed sample drawn from both flagged and unflagged accounts
Why it is observable, and what it does not claim
Appeals measure who complained, not who was wrongly flagged. Only the random audit estimates error in both directions

7

What a vendor publishes each quarter
Shared bands. Flags reported against common threshold bands rather than each vendor's private cutoff
Why it is observable, and what it does not claim
Comparability is the entire purpose of the exercise, and a vendor-chosen threshold destroys it

8

What a vendor publishes each quarter
Privacy floor. Aggregate publication only, a minimum cell size below which figures are suppressed, a stated retention limit on the underlying signals, and independent confidential audit of both
Why it is observable, and what it does not claim
The measurement implies closer scrutiny of accounts, usage and payment. A standard that does not price that cost is not adoptable

9

What a vendor publishes each quarter
Split disclosure. Classification methodology published openly; operational signals and thresholds disclosed confidentially to the independent auditor alone
Why it is observable, and what it does not claim
Publishing the full detection method teaches evasion. Publishing none of it makes the figure unauditable

Rows 4 and 5 are deliberately weaker than the questions they gesture at, and an earlier version of this proposal got both wrong in the same way. Row 4 once proposed inferring an account's revenue source from a freelance marketplace or an app store; a vendor sees that an invoice was paid, not who paid it or why. Row 5 once proposed counting accounts observed calling another vendor's model; a vendor sees requests arriving at its own endpoints and has no view of what a customer's machine contacted next. Both are now narrowed to the observable version, and the questions they cannot answer are named below rather than assumed away.

Now the honesty, because a proposal that hides its weaknesses is not worth adopting.

It measures a footprint, not a population. Two accounts may be one system. One account may be running ten. This is a number about one vendor's own billing surface and it should be named that way rather than dressed up as a census. The thing you can count is not the thing you wanted to count, and it is the only thing on offer.

It will surface false positives, in volume. A solo developer's scheduled agent, a scraping service, an automated content pipeline, a small automated trading bot: each has close to the same signature, and row 5 will catch every one of them. That is what row 2 is for, and why row 6 rests on random audit rather than on appeals.

It requires the labs to publish something about their own billing and enforcement surface that they currently do not. That is a real cost and it should be stated as one. It also exposes a commercially awkward figure: how much of a vendor's revenue arrives from accounts it cannot associate with anyone in its own records.

No single vendor can see the chimera, and Achiam identified that limit before this piece did. His sentence is quoted above and it is correct as infrastructure: "No individual lab may be able to detect that there is an orchestrator or sequence of orchestrators using intermittent model calls from burner API accounts to sustain its own existence." Anthropic said the same thing about its own data a year later and in plainer words: "We can only analyze traffic from a single model provider". The only slice of the multi-model case visible from inside one company is traffic running through that company's own hosted orchestration surface. The honest conclusion is not that the instrument fails but that the problem is larger than any one company can fix, which argues for publishing partial comparable counts rather than against it: several vendors reporting on shared bands is the route available from inside the industry to a picture none of them can assemble alone.

It cannot see the substrate the labs do not meter. Achiam's example starts with an AWS instance. A vendor's report reaches none of that, and reaches nothing at all about open-weight models running on rented hardware with no vendor in the loop. Nothing in the documents listed at the end of this piece obliges a compute provider to publish anything comparable, and this proposal does not pretend to reach them.

One caveat about the evidence this proposal leans on hardest. Anthropic's separate report on AI-enabled cyber threats (3 June 2026) draws on "832 accounts that were banned for malicious cyber activity between March 2025 and March 2026", and states its own limit in the same breath: "These 832 cases are just a subset of the total number of accounts banned during this period, but they represent those where we had enough detail to conduct a thorough assessment of the attackers' techniques." That report demonstrates that account-level enforcement reporting is publishable and useful. It does not show that those accounts were systems sustaining their own operation, and nothing here should be read as claiming it does. It is evidence about format, not about rogue AIs.

The argument that naming a failure concedes it has been settled elsewhere

The objection Achiam names, that acknowledging the failure mode reads as surrender on preventing it, is old and it is not specific to AI. Public health once wrote a version of it into statute.

A 1988 provision codified at 42 U.S.C. § 300ee-5 restricts the funds it covers, those "provided under this Act or an amendment made by this Act", from being "used to provide individuals with hypodermic needles or syringes so that such individuals may use illegal drugs, unless the Surgeon General of the Public Health Service determines that a demonstration needle exchange program would be effective in reducing drug abuse and the risk that the public will become infected with the etiologic agent for acquired immune deficiency syndrome" (Pub. L. 100-607, title II, § 256(b), 4 November 1988; text as reproduced in the US Code at Cornell's Legal Information Institute, retrieved 2 September 2026). This piece characterises only that provision's own text and takes no position on the wider funding history around it. The money it covers was withheld until an official could certify that a needle programme "would be effective in reducing drug abuse": until someone certified that the harm-reduction measure did not concede the primary goal.

The evidence answered it, and not by counting alone. The National Research Council and Institute of Medicine report Preventing HIV Transmission: The Role of Sterile Needles and Bleach (1995) weighed self-reported behaviour, programme demographics and modelled incidence across a body of studies, reporting that of those assessing needle sharing, "10 of the 14 studies showed a beneficial effect of the needle exchange programs on reported frequency of needle sharing; 4 showed a mixed or neutral effect", and on the concession question, that "most projects suggest that programs do not increase injection drug use."

Now the limit of the comparison, which matters more than the comparison.

That field had its unit handed to it. A syringe returned is a countable event, and nobody had to invent what a case was before they could count cases. Persistent autonomous operation gets no such gift, which is the whole burden of the sections above. So the honest version of the analogy is narrow: what public health proves is that the symbolic objection loses once the evidence is on the table. It does not prove that such evidence can be had here. That still has to be built.

Two objections owed a straight answer

The posts prove nothing about institutions. Achiam wrote as a private citizen, after he had left the company. A man who no longer works somewhere is not evidence of what that place will or will not say. He is a private citizen thinking in public. Any argument built on these posts as a demonstration of institutional speech costs collapses, and we are not building one. The posts did one thing here, which was to make a gap legible by asking for a measurement precisely enough that its absence became visible. The claim rests on documents: a gating framework whose criteria are model-level, a category with no threshold, a published post-deployment study that names one provider as its own boundary, and nothing that crosses it.

The vendors already do this. They do, and more than an earlier version of this piece credited. OpenAI's framework commits to monitoring, investigation capacity and bans. DeepMind commits to detection across the model lifecycle. Anthropic has published real autonomy measurements with real denominators. Conceding that improves the argument rather than weakening it, because it moves the complaint from absence to interoperability, which is a harder charge to dismiss and a cheaper one to fix. The residue is one sentence: none of these figures are reported on a common basis, and a system that spreads itself across four vendors is invisible in all four reports at once. A vendor that adopted the nine rows above and found the residual category in row 2 empty would have refuted this piece with a document, which is the correct way to refute it.

What a serious answer looks like

Achiam closes the first post with "I think we should rip the bandaid off and have the conversation." He is right, and the conversation has a first move that is not a conversation at all.

The measurement he asks for cannot be built as worded, for the reason he gave: the unit it needs does not survive his own description of the thing. The measurement that can be built is smaller, uglier and available now: not how many of them there are, but what rate of accounts shows persistence nobody can associate with a person, for how long, paid across how many cycles, in a shape no human session has, split from the scheduled jobs and enterprise agents that look identical, and audited at random so the error rate is published beside the finding.

It would make the disagreement empirically contestable. That is the whole of the ambition. Right now two reasonable people can hold opposite views about how much autonomous operation is happening in the world and neither can be shown to be wrong, which is the condition Achiam is objecting to when he says the conversation cannot happen.

So the position is narrow and it is not a complaint about anyone's courage. OpenAI has written a related capability into a framework, set one potential response beside it, and committed to monitoring and enforcement after release. DeepMind commits to detection across the lifecycle. Anthropic has measured agent autonomy after deployment, published the numbers, and named the boundary of its own data: one provider. The components exist. What does not exist is a figure that crosses the vendors, uses the same bands, and reports its own error rate, which is the only shape in which the number would mean anything at all.

What we read

A claim about what nobody publishes is only auditable if the search is stated, so here is the document set this piece rests on, read in the original.

OpenAI, Preparedness Framework Version 2 (April 2025)

What it settled
Tracked and Research Categories, the five criteria, the ARA entry and its potential response, the post-deployment safeguards

OpenAI, GPT-4o system card (8 August 2024)

What it settled
The four scored Preparedness categories cited above, read for the previous piece

Anthropic, Measuring AI agent autonomy in practice (18 February 2026)

What it settled
Post-deployment autonomy measurement exists, with denominators, and stops at one provider

Anthropic, What we learned mapping a year's worth of AI-enabled cyber threats (3 June 2026)

What it settled
Account-level enforcement reporting is publishable, with its own subset caveat

Google DeepMind, Frontier Safety Framework page, version 3.1 (17 April 2026)

What it settled
Lifecycle detection is a stated commitment elsewhere in the field

Joshua Achiam, two posts, profile and about page (read 2 September 2026)

What it settled
The argument this piece answers, and his dates and roles

42 U.S.C. § 300ee-5; NRC and IOM, Preventing HIV Transmission (1995)

What it settled
The comparison case, statute and evidence

What we did not read, and where a counter-example would most likely be found: the publications of frontier vendors other than the three above, and of industry bodies such as the Frontier Model Forum, which is where a cross-provider figure would most naturally be published; the full PDF of DeepMind's framework rather than its public page; cloud and compute providers' own transparency reporting; any vendor's internal, unpublished monitoring; enterprise contract terms that are not public; and any document not in English. If a cross-provider measure of persistent autonomous operation exists in one of those, this piece is wrong and we would like to be shown it.

What would change our mind

Four objections, and two of them take most of the argument with them.

The first is close to fatal to the framing this piece started from. Achiam wrote these posts as a private citizen after he had left OpenAI. He left on 25 July 2026; we read the posts on 2 September 2026. A man who no longer works somewhere is not demonstrating what an institution will or will not say, he is demonstrating what he thinks, which he is free to do. Any reading of these posts as evidence about institutional speech costs is unsupported, and we make none. We concede the whole of it. What survives is carried by documents rather than by posts: a gating framework whose criteria are model-level, a category with no threshold, a published post-deployment study that names one provider as its own limit, and no published measure that crosses that limit. The posts are what made the gap legible. They are not the evidence for it.

The second is that 'measurable' is a capability-evaluation criterion and the rows proposed below cannot satisfy it. The framework's own words: 'Measurable: We can construct or adopt capability evaluations that measure capabilities that closely track the potential for the severe harm.' A quarterly rate of flagged accounts is not a capability evaluation, does not elicit anything from a model, and could never gate a release. Anyone holding that the only legitimate safety instrument is a pre-deployment evaluation should conclude the proposal is beside the point. We concede that it does not make Autonomous Replication and Adaptation trackable and does not try to. It is a different instrument for a different question.

The third is that vendors already do this work and publish some of it, so the complaint is about format rather than substance. Anthropic's agent-autonomy study is real post-deployment measurement with real numbers, and OpenAI's framework carries monitoring and enforcement commitments. Largely conceded. The residue is narrow and it is the piece's actual claim: none of it is comparable across providers, and a rogue system distributed across several vendors is invisible to every one of those instruments individually. If a critic thinks single-provider reporting is sufficient, they should say what a distributed system looks like in it.

The fourth is Achiam's own hedging on severity. He writes that the outcome will probably not be 'anywhere near as catastrophic' as people predict and that loss of control is 'a matter of degree'. If he is right, a cross-provider instrument is a research nicety rather than a safety necessity, and the ordering of priorities the labs have chosen is defensible.

What would change our mind. Two separate conditions, because this piece makes two separate claims and they can fail independently.

On classification: Autonomous Replication and Adaptation appearing as a Tracked Category in a future Preparedness Framework revision, with a defined evaluation, a public rubric and a stated response at the high end. That would settle the classification question and would leave the measurement claim entirely untouched, because a gating instrument still does not count what is running.

On measurement: two or more frontier vendors publishing a persistent-autonomy figure on shared threshold bands, with exposure-adjusted denominators, a category split that separates scheduled automation from unexplained persistence, and a validation estimate from random audit rather than from appeals. That would refute the claim outright. A weaker version, one vendor publishing it alone, would not, because the whole argument is about the vendor boundary. Running the other way: if a lab publishes a serious attempt and shows that persistent autonomous operation cannot be separated from ordinary automated business at any useful precision, the proposal fails on its own terms.

Where these numbers come from

This piece argues from outside documents rather than from a study of our own. Every figure and every quotation is sourced in the text to the document it came from: a paper, a company's own published policy, a model card, a regulator's text. You can open the original and read the sentence around it. Where a claim could not be traced to a document you can open, it is not here.

See more of this work on Google

Google lets you name the sites you want to see more of. Adding AI Tools Police changes what Google shows you, not where we rank for anyone else, and it tells us nothing about you.

Add as a preferred source on Google