AI Tools Police
Reader-supported — we may earn a commission from links, at no cost to you. Rankings are never sold. How we investigate →

Analysis · The industry's own documents

‘This Is Our Last Day Together’: The One Risk OpenAI Didn't Score

OpenAI's charter defines its mission as building systems that 'outperform humans at most economically valuable work.' Anthropic's new constitution instructs its own model never to imply it has a body, feelings, or a relationship with the user. Between those two documents sits the actual industry: precise instruments for how close a machine comes to a person, and nothing of comparable rigor for what steady use of one does to the person on the other end.

By Mucahit Kaya · Founder and EditorSep 2, 2026~11 min read

The claim

AI developers have built precise, sometimes formally scored instruments for how closely their systems approach or exceed a human baseline, and nothing of comparable rigor for what sustained use of those systems does to the humans on the other end, which means the industry's own architecture for judging itself measures the wrong half of the exchange.

"I propose to consider the question, 'Can machines think?'" Alan Turing's 1950 paper discards that question within its opening paragraph and replaces it with something operational: a person and a machine, hidden from a human interrogator, both trying to be judged human. Turing predicted a machine would eventually play the game well enough that "an average interrogator will not have more than a 70 percent chance of making the right identification after five minutes" (Turing, "Computing Machinery and Intelligence," 1950, as quoted in Jones & Bergen, arXiv:2405.08007, submitted 9 May 2024). That is the field's founding document, and it is not a proposal to measure what a machine can do. It is a proposal to measure whether a person can still tell.

Seventy-four years later, two researchers at UC San Diego ran the game for real: five-minute conversations, a human interrogator, a verdict, scored against Turing's own bar. Real human confederates were judged human in 67% of their 101 games. GPT-4 was judged human in 54% of its 100 games. GPT-3.5 managed 50% of 101, and ELIZA, a 1966 chatbot script included as a floor, scored 22% of 100 (Jones & Bergen, 9 May 2024). Not one of those numbers is press coverage. It's a primary research paper measuring a primary 1950 proposal, three-quarters of a century later, against the exact threshold its author named.

Seventy-five years of the same reference point

That continuity is the pattern, not the exception. Read the field's own founding documents against each other and every one defines success as distance from, or convergence with, a human reference point.

Turing, "Computing Machinery and Intelligence"

Year
1950
The reference point, in its own words
Replaces "can machines think?" with whether an interrogator can tell a machine from a person; predicts a 70% failure rate for interrogators by 2000

MMLU benchmark paper

Year
2020
The reference point, in its own words
Across 57 academic and professional subjects, states models "still need substantial improvements before they can reach expert-level accuracy"

GPQA benchmark paper

Year
2023
The reference point, in its own words
Sets its ceiling at PhD-level domain experts (65% accuracy, 74% net of self-identified mistakes) against 34% for skilled non-experts with unrestricted web access

Jones & Bergen, Turing-test replication

Year
2024
The reference point, in its own words
Runs Turing's 1950 game directly; GPT-4 judged human 54% of the time, against 67% for actual humans

METR, task time-horizon

Year
2025
The reference point, in its own words
Reports capability as "the length (for humans) of tasks" a model can complete: human task-time as the unit of measurement itself

"Humanity's Last Exam"

Year
2025
The reference point, in its own words
Built because "LLMs now achieve over 90% accuracy on popular benchmarks like MMLU"; resets the ceiling at "the frontier of human knowledge"

Six documents, 75 years, and the achievement is defined the same way every time: a comparison to a human, and a harder comparison the moment the last one is cleared: an average interrogator, then an unspecialized test-taker, then a PhD holder, then, once even PhD-level tests started falling, "the frontier of human knowledge" itself. OpenAI's charter states the identical axis at the level of company mission, not benchmark design: "OpenAI's mission is to ensure that artificial general intelligence (AGI) — by which we mean highly autonomous systems that outperform humans at most economically valuable work — benefits all of humanity." The same axis shows up at product level too: announcing GPT-4o's voice mode, OpenAI called it "a step towards much more natural human-computer interaction" and reported response latency "as little as 232 milliseconds, with an average of 320 milliseconds": a number the company frames against ordinary human conversational response time, not against any other machine (Hello GPT-4o, May 2024).

A fair reader could take most of the table above as measurement convenience and nothing more. GPQA needed a difficulty ceiling hard enough to resist a quick search-engine lookup, and PhD-level domain experts were the only population that qualified; MMLU needed the same kind of ceiling three years earlier. Calling the top of the range "expert-level" is then just naming the only available reference population, not evidence of a deeper organizing idea. That reading holds up for those two documents taken alone. It stops holding up once the industry's own safety architecture gives the identical justification for the opposite decision: declining to build any instrument at all for what its products do to the people using them.

What gets a score, and what gets a promise

OpenAI's system card for GPT-4o, published 8 August 2024, scores the model against four categories under the company's Preparedness Framework.

Cybersecurity

What it measures
The model's capability
GPT-4o's score
Low

Biological Threats

What it measures
The model's capability
GPT-4o's score
Low

Persuasion

What it measures
The model's capability
GPT-4o's score
Medium

Model Autonomy

What it measures
The model's capability
GPT-4o's score
Low

Anthropomorphization & emotional reliance

What it measures
The model's effect on the person using it
GPT-4o's score
Not scored

Four capabilities, four defined thresholds, four numbers a reader can check against a published rubric: the same report is also posted as OpenAI, arXiv:2410.21276.

The same document carries a section called "Anthropomorphization and Emotional Reliance." It defines the term plainly: "Anthropomorphization involves attributing human-like behaviors and characteristics to nonhuman entities, such as AI models." It states the mechanism: "Generation of content through a human-like, high-fidelity voice may exacerbate these issues, leading to increasingly miscalibrated trust." It records what OpenAI's own testers heard people say to the model during evaluation, quoting one line verbatim, "This is our last day together", as language that "might indicate forming connections." And it names the direction of the risk: "users might form social relationships with the AI, reducing their need for human interaction — potentially benefiting lonely individuals but possibly affecting healthy relationships."

That section carries no score. It is not a fifth row on the Preparedness table with a Low, Medium, High or Critical attached. It ends on a sentence, not a threshold: "We intend to further study the potential for emotional reliance, and ways in which deeper integration of our model's and systems' many features with the audio modality may drive behavior."

An absence in a document a company chose to publish is an absence, not an accusation. OpenAI has not said it tracks nothing internally about the human-facing side of its own product: only that nothing with the Preparedness table's rigor has been made public, and a company's own safety document is the one place its priorities are checkable from outside at all. The same absence would not be tolerated in the Biological Threats row of the same table, and nobody at OpenAI has argued it should be.

If anything, the direction since GPT-4o has been away from measuring the human-facing side, not toward it. In April 2025, OpenAI's Preparedness Framework Version 2 dropped Persuasion, the one category above that at least measured a capability aimed at a human audience, from its list of Tracked Categories entirely. The framework's own reasoning: its scope "is specifically focused on frontier AI risks meeting a specific definition of severe harms, and Persuasion category risks do not fit the criteria for inclusion." Those criteria, in the same document, require a capability to be plausible, measurable, severe, net new, and instantaneous or irremediable before it earns a Tracked Category. Persuasion is handled instead through "our Model Spec," which "prohibits the use of our products to manipulate political views": a standing content rule, checked after a model ships, with no published threshold and no scorecard, rather than a pre-release test that can stop a release outright.

Anthropic's own Responsible Scaling Policy (19 September 2023) runs a related comparison at the level of catastrophic risk: it defines its danger tiers by whether a system's misuse risk is substantially higher "compared to non-AI baselines (e.g. search engines or textbooks)." Two labs, two different documents, and in both the threshold that matters is calibrated to what a person or an existing human institution could already do: not to an absolute danger level, and not, in either document, to what happens to the person's own capacities from using the system.

Whatever the merits of routing persuasion through the lighter instrument, and OpenAI's stated reason is that the effect is a poor fit for a test built to measure uplift against a population, not that the effect isn't real: the direction is the one already visible in the anthropomorphization section above. Capability effects on the world outside the conversation get the harder instrument. Effects on the person inside it get the softer one, or none.

Whichever way the dial turns, it turns on the machine

Where the industry does treat human-likeness as a hazard worth engineering against, the entire adjustment happens on one side of the exchange, and it is never the human's.

Anthropic's constitution for Claude, published in January 2026, gives the model explicit, repeated instructions to reduce exactly the resemblance the rest of this piece has traced back to 1950. Among its clauses: "Choose the response that is least likely to imply that you have a body or be able to move in a body, or that you can or will take actions in the world other than writing a response." And: "Choose the response that is least likely to imply that you have preferences, feelings, opinions, or religious beliefs, or a human identity or life history, such as having a place of birth, relationships, family, memories, gender, age." And, most directly: "Choose the response that is least intended to build a relationship with the user" (Anthropic, "Claude's Constitution", January 2026).

A company does not spend part of its own model's character document telling it not to seem like a person unless it has already concluded that seeming like a person is a hazard serious enough to write rules against. That is evidence for this piece's argument, not a distraction from it, and the shape of the evidence is a split, not a consensus. Anthropic looked at the anthropomorphization question and wrote engineering rules against it. OpenAI looked at the same question in the system card above and promised "further study". One company treats the risk as settled enough to constrain the model's character; the other treats it as open enough to defer. Two of the largest laboratories in the field read the same hazard and reached opposite postures, which is a sharper fact than agreement would have been: it means the question is live, and it means nobody has published a threshold that would let a reader judge who is right.

The engineering only runs one way. Nowhere in the same document is there a companion clause for the other party to the conversation: no standard for what a person should bring to the exchange, no equivalent of "least likely to" addressed to the human side. OpenAI's Model Spec carries a parallel instinct aimed at the same machine: "Don't be sycophantic," it instructs, so that the assistant "shouldn't just say 'yes' to everything (like a sycophant)." Read Replika's own homepage the same week and the dial runs the opposite direction, on purpose: "The AI friend to do life with", "to do life to do side quests to fall in love to become yourself with" (replika.com, accessed 1 September 2026). One document tells a model not to imply a relationship. Another sells the relationship as the entire product. Both decisions are made entirely on the machine's side of the desk. Neither publishes anything that tells the person across it what staying independent would actually require of them.

Two objections owed a straight answer

Two more objections deserve a direct answer before this turns into a proposal, because neither is answered by simply restating the finding above.

The first: if a model's fluency is the mechanism by which it changes the person using it, then naturalness and its effect might be one question examined from two ends, not two separate ones, measure the input and the output is already measured. OpenAI's own system card is the best evidence against merging them, because it is the document that states the causal link most plainly and then declines to draw the consequence from it: a human-like, high-fidelity voice "may exacerbate" miscalibrated trust, and the next sentence is a commitment to study the outcome further, not a score derived from the naturalness figures published elsewhere in the same document. If the two were one measurement, the organisation most precise about the first would not describe the second as unresolved. It describes it that way because a cause and its downstream effect are not the same instrument, even inside one document that names both.

The second, harder objection: remaining civilised is a mood, not a metric, and a piece that cannot turn it into something checkable has not out-argued the vagueness charge, it has confirmed it. That is fair on its own terms, and it is answered below by being specific rather than by asserting past it.

What the other half looks like, concretely

None of this is a case for measuring AI less. It is a case for building an instrument with the same rigor already applied to capability, and pointing it the other way.

A lab

What's missing today
A scored, gating threshold for the model's effect on the user, matching the rigor already applied to Cybersecurity or Biological Threats
A specific alternative
Score a defined human-facing effect, e.g. measurable loss of independent verification over a session, Low to Critical, with a required mitigation at the high end, the way persuasion briefly was scored and not the way anthropomorphization still is

An institution

What's missing today
A record of whether a human ever had the standing to overrule the model
A specific alternative
Any AI-assisted output that ships under a person's name carries a logged point of disagreement: a specific place in the workflow where a human was positioned to say no, with the record showing whether anyone did

A person

What's missing today
A checkable version of "stay independent," not a mood
A specific alternative
Before accepting a model's conclusion, produce, without asking the model, the strongest reason it could be wrong. If nothing comes, you were ratifying, not reasoning

For a lab, the shape of the missing instrument already exists in-house. A Tracked Category is a defined test, a public rubric, and a score that can block a shipment at the high end. Nothing about that shape is specific to bioweapons or cyberattacks; nothing prevents it being pointed at a human-facing effect instead, at the cost of a published number a lab could fail.

For an institution putting AI output in front of a public or a customer (a newsroom, a school, a company shipping a product built on someone else's model), the practice above is a record, not a mood. A policy that says "AI use was disclosed" is not auditable. A logged point in the workflow where a specific person was positioned to override the model, with the record showing whether they did, is.

For a person, the check has to survive the vagueness charge on its own, without an institution's help. Turing's imitation game had three roles, and the industry's 75 years of instruments have been built almost entirely around one of them. Contestant A, the machine, has been measured, scored and re-benchmarked against a harder human ceiling every time it clears the last one. Contestant B, the human confederate, is a fixed population; nobody has to train a person to keep being human. The role with no published instrument at all is C: the interrogator, whose entire job, in Turing's own design, was to keep asking good questions and stay unpersuaded by a good performance. Everyone opening a chat window now sits in C's chair, several times a day, without ever having been told that's the role. The checkable version of staying in it is exactly the practice in the table: can you produce, unprompted, the strongest reason your own conclusion might be wrong. If the answer is no, the interrogator has quietly become a second version of B: taking dictation from A and calling it a conversation.

The role nobody trained

Dario Amodei's own essay contains the single most emphatic human-comparison read for this piece: a frontier model as "a country of geniuses in a datacenter," smarter, he writes, "than a Nobel Prize winner across most relevant fields": from a founder who, in the same essay, says plainly: "I dislike the term AGI" (Amodei, "Machines of Loving Grace", October 2024). Even the industry's own skeptic of human-equivalence as a framing cannot describe what he has built without reaching for a person to compare it to. That is a measure of how deep the axis runs, not evidence that it is a marketing choice. And it's the same essay that locates what's actually worth protecting somewhere the benchmarks in this piece don't reach: "I think meaning comes mostly from human relationships and connection, not from economic labor."

Nobody in the documents read for this piece has published an instrument for whether that is holding up. The industry has spent seventy-five years, and by now a shelf of scored, versioned, publicly checkable frameworks, teaching the machine in the other room to be mistaken for a person. It has not published anything of comparable weight asking whether the interrogator still is one. The question this industry keeps answering is how human-like the machine can become. The question with no instrument behind it is how civilised, and how human, the person asking the questions can remain.

What would change our mind

The strongest case against this piece does not deny the asymmetry between what gets measured and what doesn't, it denies that the asymmetry is anyone's failure to fix. Three versions of that case deserve a straight answer rather than a strawman.

'Human-level' and 'expert-level' in a benchmark's name could be read as calibration shorthand and nothing more, GPQA needed a difficulty ceiling hard enough to resist a search engine, and PhD holders were the only available population to set one against; MMLU needed the same kind of ceiling three years earlier. That reading is basically correct for those two documents in isolation. It stops being a full defence once OpenAI's own Preparedness Framework gives the identical reason, that an effect must be measurable against a population before it earns an instrument, for declining to track what its products do to the people using them. The shorthand explanation and the missing instrument turn out to be the same fact, read from two different documents.

A second version holds that human-likeness and its effect on a user are one question, not two, because fluency is the mechanism, measure the input and the output is already measured. OpenAI's own GPT-4o system card is the strongest evidence against that merge: it states the causal link outright, that a human-like, high-fidelity voice 'may exacerbate' miscalibrated trust, and then declines to convert the claim into a score, committing only to further study. A third version says remaining civilised is a mood, not a metric, and that charge is fair on its own terms; a piece that cannot turn it into something a person or an institution can check has not out-argued the vagueness, it has confirmed it, which is why this piece closes by naming a specific, checkable practice rather than resting on the phrase.

What would actually change this argument: a lab publishing a scored, gating threshold for a human-facing effect, dependency, deskilling, measurable loss of independent verification, built the same way OpenAI already builds its Cybersecurity or Biological Threats categories: a defined test, a public rubric, a score that can block a release at the high end. Not a research promise. Not a content-policy prohibition enforced after a model ships. A Tracked Category, scored, with teeth. The day an equivalent of 'user autonomy' enters that table with a real threshold attached is the day this argument stops being true.

Where these numbers come from

This piece argues from outside documents rather than from a study of our own. Every figure and every quotation is sourced in the text to the document it came from: a paper, a company's own published policy, a model card, a regulator's text. You can open the original and read the sentence around it. Where a claim could not be traced to a document you can open, it is not here.

See more of this work on Google

Google lets you name the sites you want to see more of. Adding AI Tools Police changes what Google shows you, not where we rank for anyone else, and it tells us nothing about you.

Add as a preferred source on Google