Frontier model comparison · Anthropic and OpenAI · two rate cards, Anthropic's benchmark table and system card, six outside evaluators and 23 dated user reports
Claude Opus 5.5 vs GPT-6 Astra: 40% of the list price, but dearer per task at max effort
Claude Opus 5.5 lists at $4 and $20 per million tokens against GPT-6 Astra's $10 and $50, and its cache reads cost a fifth of Astra's. Per task the order depends on the effort setting: Artificial Analysis measures Opus 5.5 dearer at max and xhigh, cheaper at medium, and higher-scoring at every setting but low. Both rate cards, Anthropic's benchmark table and system card, six outside evaluators and 23 dated user reports, read on 23 September 2026.
Our verdict
Send coding agents and knowledge work to Claude Opus 5.5 at medium or high effort, where Artificial Analysis scores it above GPT-6 Astra at the same setting for a similar or lower cost per task; keep GPT-6 Astra for jobs that need max effort at the lowest cost per task and for cross-app business automation, the two places the outside numbers favour it.
Claude Opus 5.5 vs GPT-6 Astra specs
- Released
- Claude Opus 5.5: 22 September 2026, per Anthropic's announcement. GPT-6 Astra: 3 September 2026, per OpenAI's announcement
- API model ID
- claude-opus-5-5 (Anthropic's models overview) / gpt-6-astra (OpenAI's model page)
- Context window
- Claude Opus 5.5: 1M tokens. GPT-6 Astra: 1,050,000 tokens. Both read 23 September 2026
- Maximum output
- Claude Opus 5.5: 128K tokens. GPT-6 Astra: 128,000 tokens
- Knowledge cutoff
- Claude Opus 5.5: June 2026, stated as its reliable knowledge cutoff. GPT-6 Astra: 30 April 2026
- Effort settings and default
- Both: low, medium, high, xhigh and max. Claude Opus 5.5 defaults to medium. OpenAI's model page, read 23 September 2026, does not state Astra's default
- Thinking
- Claude Opus 5.5: adaptive and always on, cannot be disabled in the API. GPT-6 Astra: reasoning token support, per the model page
- Input and output
- Both: text and image input, text output
Claude Opus 5.5 vs GPT-6 Astra pricing
- Input, per 1M tokens
- Claude Opus 5.5 $4.00 / GPT-6 Astra $10.00. Opus 5.5 is 40% of Astra (our division)
- Output, per 1M tokens
- Claude Opus 5.5 $20.00 / GPT-6 Astra $50.00. Opus 5.5 is 40% of Astra (our division)
- Cache read, per 1M tokens
- Claude Opus 5.5 $0.20 (5-minute TTL) / GPT-6 Astra $1.00 cached input. Astra is 5x (our division)
- Cache write, per 1M tokens
- Claude Opus 5.5 $5.00 / GPT-6 Astra $12.50. Both are 1.25x the vendor's own input rate
- Long prompts
- Claude Opus 5.5: the pricing page we read prints no long-context surcharge. GPT-6 Astra: prompts over 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request, which is $20 input, $2.00 cached input and $75 output (our arithmetic)
- Batch
- Claude Opus 5.5: 50% saving, so $2.00 and $10.00 (our arithmetic). GPT-6 Astra: Batch and Flex at 50% of Standard, so $5.00 and $25.00 (our arithmetic)
- Fast mode
- Claude Opus 5.5: $8.00 and $40.00, 2x standard, stated as up to 2.5x speed. GPT-6 Astra: 2x the applicable rates, so $20.00 and $100.00 at short context (our arithmetic)
- Regional processing
- Claude Opus 5.5: US-only inference at 1.1x input and output. GPT-6 Astra: OpenAI's pricing page, read 5 September 2026, charged a 10% uplift on data-residency endpoints for eligible models released on or after 5 March 2026; Astra's eligibility not checked
Rate card read at the vendor’s own documentation on Sep 23, 2026.
Claude Opus 5.5 lists at $4 per million input tokens and $20 per million output, 40% of GPT-6 Astra's $10 and $50. Which one costs less per task depends on the effort setting. On Artificial Analysis's Intelligence Index v4.3.2, Opus 5.5 costs more per task than Astra at max effort ($5.98 against $3.26) and less at medium ($1.34 against $1.54), and it scores higher at both.
The benchmark tables need the same care. Anthropic's launch table sets Opus 5.5 at xhigh against Astra at high on Terminal-Bench 4.0, each lab's highest reported score, and every Astra figure in it comes from OpenAI or a third party, never from a run by Anthropic. On Artificial Analysis's own harness the two score the same, 59.6%.
We ran neither model, and every figure below names its publisher and setting. Anthropic's documents, OpenAI's model page for Astra and six outside evaluators were read on 23 September 2026; OpenAI's full rate card and safety overview are from our reading of 3 and 5 September; the 23 user comments were posted on 22 and 23 September. Each model's full story is in its own review, and this page covers where the two differ.
Which is better, Claude Opus 5.5 or GPT-6 Astra?
For coding agents and knowledge work run at medium or high effort, Claude Opus 5.5 is the better value on the published evidence. At those settings Artificial Analysis scores it above GPT-6 Astra (51.24 against 49.57 at medium, 53.58 against 50.92 at high) for a similar cost per index task ($1.34 against $1.54, and $1.82 against $1.73). GPT-6 Astra is the better pick when a job needs max effort and cost per task decides it, and for business automation of the kind Zapier's AutomationBench measures.
Three facts drive the split. Opus 5.5 costs 40% of Astra's list price and a fifth of its cache-read price, and Astra reprices a whole request once input passes 272,000 tokens. But at max effort Artificial Analysis measured Opus 5.5 writing about 119,000 output tokens per index task against about 27,000 for Astra, so at that setting the lower rate buys fewer tasks.
The split rests mostly on one publisher's composite of ten evaluations, which can hide the single evaluation your workload depends on. Where a benchmark below matches your work, that row should outweigh our summary.
Is Claude Opus 5.5 cheaper than GPT-6 Astra?
Per token, yes. Claude Opus 5.5 costs 40% of GPT-6 Astra's price on input and output ($4 and $20 against $10 and $50 per million tokens), 40% on cache writes ($5 against $12.50) and one fifth on cache reads ($0.20 against $1.00), our division of both vendors' rates as read on 23 September 2026. The full cards, with every mode and multiplier, are in the Opus 5.5 rate card in our review and the Astra rate card in our review.
The 40% survives the pricing modes. Both vendors halve their rates for batch work and double them for fast mode, so Opus 5.5 stays at 40% of Astra in each: $2 and $10 against $5 and $25 in batch, $8 and $40 against $20 and $100 in fast mode, the Astra figures and the batch figures being our arithmetic from each vendor's stated multiplier. Regional processing: Anthropic charges 1.1x for US-only inference, and OpenAI's pricing page, read 5 September, charged a 10% uplift on data-residency endpoints for eligible models, a list we have not checked Astra against.
The cache-read ratio is the one that moves real bills. Anthropic says cache reads "make up the majority of agentic and coding work costs", and both vendors set cache writes at 1.25 times their own input rate, so the write gap stays at 40% while the read gap widens to a fifth. An agent that resends a long cached context every turn pays Astra five times what it pays Opus 5.5 for that context. The rate is still only half the bill; the other half is how many tokens each model spends on a task, covered below.
What happens to GPT-6 Astra pricing above 272K tokens?
Above 272,000 input tokens, GPT-6 Astra bills the entire request, the first 272,000 tokens included, at double its input and cache rates and one and a half times its output rate. Head to head, that cuts Claude Opus 5.5's share of the bill from about 35% to about 20% on a mostly cached prompt, because the Anthropic pricing page we read prints no long-context surcharge line. The exact wording of OpenAI's rule is quoted in our GPT-6 Astra review.
Three worked examples show where the line falls. Each carries 20,000 new input tokens with the rest of the prompt read from cache, and 10,000 output tokens.
| Prompt size | Claude Opus 5.5 | GPT-6 Astra | Opus 5.5 as a share of Astra |
|---|---|---|---|
| 271,000 tokens | $0.330 | $0.951 | 35% |
| 273,000 tokens | $0.331 | $1.656 | 20% |
| 300,000 tokens | $0.336 | $1.71 | 20% |
Our arithmetic from each vendor's published rates, read 23 September 2026. It assumes the cache was written by an earlier call and is still live, and leaves out cache writes and tool fees. Opus 5.5 is priced at list with no long-context surcharge, which is what the page we read shows.
Adding 2,000 tokens to the prompt raises Astra's bill by 74% and Opus 5.5's by less than a tenth of a cent, and the dollar gap between the two more than doubles, from $0.621 to $1.325, all our arithmetic. Below the line Opus 5.5 costs about a third of Astra on this kind of request; above it, about a fifth. On our reading of OpenAI's fast-mode multiplier, which applies to "the applicable rates", fast mode stacks on top and prices Astra's output over the line at $150 per million.
One caveat bounds all of this. The Claude pricing page we read links to a detailed pricing page we did not open, and Opus 5.5 also has a 1M-token window. If a long-context rule for Opus 5.5 sits on that page, the 20% rises. For prompts that carry whole repositories or document sets, count input tokens before a call goes to Astra.
Which costs less per task, Claude Opus 5.5 or GPT-6 Astra?
It depends on the effort setting. On Artificial Analysis's Intelligence Index v4.3.2, Claude Opus 5.5 costs more per index task than GPT-6 Astra at max ($5.98 against $3.26) and at xhigh ($3.46 against $2.31), about the same at high ($1.82 against $1.73), and less at medium ($1.34 against $1.54) and low ($0.55 against $0.82). Opus 5.5 scores higher at every setting except low.
| Effort setting | Opus 5.5 index score | Opus 5.5 cost per index task | Astra index score | Astra cost per index task |
|---|---|---|---|---|
| max | 57.62 | $5.98 | 52.67 | $3.26 |
| xhigh | 55.99 | $3.46 | 52.39 | $2.31 |
| high | 53.58 | $1.82 | 50.92 | $1.73 |
| medium | 51.24 | $1.34 | 49.57 | $1.54 |
| low | 42.31 | $0.55 | 45.78 | $0.82 |
Artificial Analysis Intelligence Index v4.3.2, read 23 September 2026. Unrounded scores and the Astra rows come from the data behind the charts on the publisher's Opus 5.5 page, a v4.3.2 page, which displays Opus 5.5 as 58, 56, 54, 51 and 42. Opus 5.5 entries run with Anthropic's default fallback enabled.
Token count explains the order. The publisher's launch article: "Opus 5.5 (max) uses ~119k output tokens per Intelligence Index task, against ~73k for Opus 5 (max), ~78k for Fable 5.1 (max) and ~27k for GPT-6 Astra (max)". At max, Opus 5.5 writes about 4.4 times Astra's output at 40% of Astra's output rate, our arithmetic, and the tokens win. The cost figures also fold in each model's typical cache-hit rate, which the publisher says it measures live.
Which row applies depends on how you run the model. Anthropic gives Opus 5.5's default effort as medium, the row where it is both cheaper and higher-scoring. OpenAI's model page names the same five settings for Astra, "reasoning.effort supports low, medium, high, xhigh, and max.", and states no default.
Two cautions. Our Astra review of 6 September printed Astra at 55 at max on index v4.2, a different version that cannot be set against 52.67. And cost per task varies by benchmark: on Zapier's AutomationBench at max, Opus 5.5 costs $1.28 per task and Astra $1.73.
Is Claude Opus 5.5 really a fifth of GPT-6 Astra's cost?
That is Anthropic's claim for particular benchmarks at particular settings, and no outside cost figure confirms it; on Artificial Analysis's full index the nearest pair is 41%. On that index, Claude Opus 5.5 at medium scores 51.24 against GPT-6 Astra at max on 52.67, at 41% of Astra's cost per task, our division. Anthropic's GDPval-AA line, "At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task.", holds on score in the publisher's data, 1576 against 1542 Elo.
| What Anthropic says | Settings it names | What the outside data shows |
|---|---|---|
| Beats Astra on GDPval-AA for about a fifth of the cost per task | Opus 5.5 medium, Astra max | Score holds, 1576 against 1542 Elo. No GDPval-only cost is published |
| Matches Astra on Terminal-Bench 4.0 at about 40% of the cost | None | Level at 59.6% only with Opus 5.5 at xhigh or max; at medium it scores 52.5% |
| Beats Astra's best FrontierCode score at roughly 20% of the cost per task | Opus 5.5 medium | Cognition's result files carry token counts and no cost, per Anthropic's card, so the ratio is Anthropic's calculation |
Anthropic's announcement and system card (p. 176); Artificial Analysis's page data for both models. All read 23 September 2026.
The Terminal-Bench line, "It matches GPT-6 Astra at about 40% of the cost.", names no setting. The match needs Opus 5.5 at xhigh or max, where Artificial Analysis's composite cost per task favours Astra, and no Terminal-Bench-only cost is published to settle the 40%.
Claude Opus 5.5 vs GPT-6 Astra benchmarks
On Artificial Analysis's own Terminal-Bench 4.0 run, Claude Opus 5.5 and GPT-6 Astra are level at 59.6%, while Anthropic's launch table shows 66.4% against 57.9%. The launch table pairs Anthropic's run of Opus 5.5 at xhigh with OpenAI's own report of Astra at high; Artificial Analysis ran both on one harness, mini-swe-agent, with three repeats over 66 tasks. Across the other benchmarks where one publisher has both, Opus 5.5 leads on ARC-AGI-2 at high, Artificial Analysis's index, GDPval-AA and Humanity's Last Exam, and Astra leads on Zapier's AutomationBench and on ARC-AGI-2 at max.
The publisher's words are that Opus 5.5 "scores 59.6%, level with the leader GPT-6 Astra (xhigh)", Opus 5.5 reaching it at both max and xhigh. Anthropic's footnote gives its own row's settings: "Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI". The card says Astra sits at high because "in their reporting max effort fared slightly lower than high, so we picked their highest number" (p. 178), and Opus 5.5 at max scored 64.8%, within noise of xhigh. So the 8.5-point gap, our subtraction, compares each lab's best reported number from separate runs, Anthropic's in Claude Code and OpenAI's in a setup the card does not describe, and Anthropic does not present it as a matched run. Its run also had safeguards on, with a fallback model answering 2.5% of requests and touching 10% of trials.
| Benchmark | Setting | Who ran it | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|---|
| Terminal-Bench 4.0 | Opus 5.5 max and xhigh, Astra xhigh | Artificial Analysis | 59.6% | 59.6% |
| Terminal-Bench 4.0 | Opus 5.5 xhigh, Astra high | Anthropic; OpenAI's report for Astra | 66.4% | 57.9% |
| Intelligence Index v4.3.2 | max, both | Artificial Analysis | 57.62 | 52.67 |
| GDPval-AA v2.1 | max, both | Artificial Analysis | 1846 Elo | 1542 Elo |
| Humanity's Last Exam | max, both | Artificial Analysis | 61.4% | 54.7% |
| AutomationBench v1.0.6 | max, both | Zapier | 40.0% | 41.4% |
| Terminal-Bench-Science 0.1 | max, both | Anthropic; OpenAI's report for Astra | 58.7% | 64.6% |
| FrontierCode v1.1 (Main) | each at its best; Opus 5.5 at medium | Cognition | 54.6% | 53.3% |
| FrontierSWE v2 | max, both | Proximal, via Anthropic's card | 62.3% | 65.5% |
| ARC-AGI-1, Semi-Private | high, both | ARC Prize | 98.5% | 98.5% |
| ARC-AGI-2, Semi-Private | high / max | ARC Prize | 93.3% / 91.7% | 92.1% / 95.0% |
Independent rows read at source on 23 September 2026, some values from the data behind the publishers' pages; vendor rows from Anthropic's announcement and system card (pp. 174 to 179). OpenAI's two figures are not checked against OpenAI's own documents.
None of the Astra cells in Anthropic's table is a run by Anthropic. The card's caption says competitor figures come from each developer's published system cards or benchmark leaderboards, and for FrontierCode it says "Cognition ran the evaluation for every model shown and reported the results", with Claude models in Claude Code and GPT models in Codex CLI. The same table prints Humanity's Last Exam with tools at 67.7% against 57.2%, with no effort stated for Opus 5.5 and no specific source for the Astra cell; the Artificial Analysis row above is the like-for-like one.
Is Claude Opus 5.5 on ARC-AGI-3 or LMArena?
Not yet. As of 23 September 2026 ARC Prize lists no ARC-AGI-3 result for Claude Opus 5.5 at any setting, so GPT-6 Astra's ARC-AGI-3 scores have nothing to be compared against. LMArena lists Opus 5.5 on none of its boards, so Astra's first place on the WebDev board is a lead over other models and not over Opus 5.5. Epoch AI's Capability Index and Artificial Analysis's Coding Agent Index likewise carry Astra and not Opus 5.5. Astra's own results on those boards are in our Astra review, linked above. An empty row one day after launch is no evidence against the model.
Where does GPT-6 Astra beat Claude Opus 5.5?
GPT-6 Astra is ahead on five published measures: Zapier's AutomationBench at max (41.4% against 40.0%), Terminal-Bench-Science in Anthropic's own table (64.6% against 58.7%, Astra's figure as reported by OpenAI), Proximal's FrontierSWE v2 at max (65.5% against 62.3%), ARC-AGI-2 at max (95.0% against 91.7%) and Artificial Analysis's index at low effort (45.78 against 42.31). It also writes far fewer output tokens per task at max effort.
Business automation across apps is the clearest case against the cheaper model. Zapier scores each task's final state against fixed criteria, and its 1.4-point gap comes with no published error margin. Anthropic says Zapier ran Opus 5.5 "without fallback models, so safeguard interventions were considered failures", which it says lowered the score; Zapier's page does not say. Opus 5.5 at max does lead Zapier's Sales domain at 42.74% and ties Astra in Operations at 62.0%.
On Terminal-Bench-Science Astra's lead is 5.9 points, our subtraction, against a standard error the card gives as ±4.8 points for Opus 5.5. At low effort Astra scores 45.78 against 42.31 on Artificial Analysis's index for 27 cents more per task. And at max, where Opus 5.5 writes about 4.4 times as many output tokens, Astra's cost per index task is 55% of Opus 5.5's, our division, for a score about five points lower.
Which is better for coding, Claude Opus 5.5 or GPT-6 Astra?
Neither leads every coding measure. Claude Opus 5.5 leads Cognition's FrontierCode (54.6% at medium against Astra's best of 53.3%) and ties GPT-6 Astra at 59.6% on Artificial Analysis's Terminal-Bench 4.0 run; Astra leads Terminal-Bench-Science and Proximal's FrontierSWE v2. At 40% of the list price, a tie or a narrow gap favours Opus 5.5 on cost, except at xhigh and max effort, where Artificial Analysis's per-task figures favour Astra.
The results divide by kind of work. FrontierCode grades whether an agent's change would be merged without edits and penalises out-of-scope changes, per the card, and Opus 5.5 scores best there at medium, its default. FrontierSWE v2 and Terminal-Bench-Science are longer agentic tasks, and Astra leads both. The first-day user reports lean further to Opus 5.5, as the user reports below show, and none of them replaces running both models on your own repository.
Claude Opus 5.5 vs GPT-6 Astra context window
GPT-6 Astra has a 1,050,000-token context window and Claude Opus 5.5 a 1M-token window; Astra caps output at 128,000 tokens and Anthropic lists Opus 5.5 at 128K. The practical difference is price: Astra bills at its long-context rates once input passes 272,000 tokens, and the Anthropic pricing page we read prints no such threshold.
Both take text and image input, return text and expose the same five effort settings. Anthropic gives Opus 5.5's default effort as medium and its reliable knowledge cutoff as June 2026, with thinking that cannot be switched off in the API; OpenAI gives Astra's cutoff as 30 April 2026 and no default effort. On long-document reasoning, Artificial Analysis's AA-LCR has Opus 5.5 at 84.7% at max in its page data and no Astra figure in our record, and the launch article says Opus 5.5 "remains behind on CritPt, AA-LCR, and GDP.pdf" without naming the leader.
What do users say about Claude Opus 5.5 vs GPT-6 Astra?
Most first-day reports that set the two side by side favour Claude Opus 5.5 on coding and on subscription quota, and they disagree about token use. We link 23 comments from 22 accounts on Reddit and Hacker News, posted on 22 and 23 September 2026. None describes GPT-6 Astra winning a concrete coding task, and one reports Astra using a third of Opus 5.5's output tokens. These are self-selected comments from one launch day, mostly about subscription plans, and they show where the argument sits without settling it.
Three describe moving back to Claude after a spell on Astra, and two more lean that way. Frumbleabumb, 22 September 2026: "But opus 5.5 is great. Token usage is low, and I get the typical high context reasoning you expect out of Claude I just can't get with Astra." ArcaneMoose, the same day, had let a two-year Claude subscription lapse for Astra a day before the launch. impulser_, 22 September 2026: "Usage is actually Claude now because of Opus 5.5". schipperai, 22 September 2026, found Astra "wasn't as good of an upgrade from Sol when it comes to coding and orchestration" and called Opus 5.5 "Fable-level performance that's cheaper and faster, and communicates plainly and briefly." qkamikaze, 22 September 2026: "I feel like Opus 5.5 is almost comparable to Astra, for a fraction of Astra's prices."
Most of the reports with a concrete coding task favour Opus 5.5. Soft-Bus-5880, 22 September 2026: "I spent hours wrestling with Astra without getting anywhere, and Opus just nailed the entire thing in a single prompt". Reasonable-Sign8458, 22 September 2026, who pays for both, re-ran code Astra had written: "Ran it with Opus 5.5 today and it fixed 98 bugs." A later comment on 23 September put the run at "3 hours and 38 minutes". roflc0ptic, 23 September 2026, back-testing against pull-request comments: "About $24 for 10 PRs. I was surprised how much worse Astra did on correctness; I stopped testing with it." senko, 22 September 2026: "Opus 5.5 is neck-and-neck with Fable 5.1 and Astra 6 in my vibe-coding tests". dondiegorivera, 22 September 2026, who uses Astra daily, wrote that "the new Opus feels like a phase shift." One found both wanting: TheVoyant, 22 September 2026, had Astra fail and revert three tasks at 64% of a $100 plan's usage, while Opus 5.5, at 43% and hitting its five-hour limit almost at once, "has done the task wrong three times."
Users split on cost per task, and two of them read the same Artificial Analysis table. Wsz2020, 22 September 2026, reproduced rows including "Claude Opus 5.5 (high) | 53.6 | $1.82" and "GPT-6 Astra (high) | 50.9 | $1.73", which match the figures above. PrinceRufusFastcar, 22 September 2026: "Certainly Opus 5.5 at 'max effort' is very expensive, but Opus 5.5 at 'high' is both a lot cheaper than Astra and scoring higher." The score half matches the data; the cost half holds against Astra at xhigh or max, and not against Astra at high, where the publisher has $1.82 against $1.73. Roland31415, 23 September 2026, wrote "30% cheaper than Astra. 75% cheaper than Opus 5.", a figure we could not match to any same-setting pair in that table; the nearest, low effort, is 33% cheaper and scores lower. ShadyShroomz, 22 September 2026, found Opus "lower cost per task in my personal benchmark, and better performance", while granting Astra a slight edge at the top tier.
On quota and tokens the reports split. this_user, 22 September 2026: "Astra is barely usable even on the $100 plan." and "Opus is at least actually usable even on the small plan." kuatecno, 23 September 2026: "Astra was wasting tokens and not following my goals.", though "I might be back with astra in a couple days". Against both, user43928, 22 September 2026, wrote that Astra "uses a third of the output tokens of Opus 5.5, Fable 5.1, or Opus 5. 17k for Astra xhigh vs 61-66k." Artificial Analysis's counts point the same way, about 27,000 against about 119,000 at max. A subscription quota and an API token count are different measures, and these comments mix the two.
Some drew the line by workload. ptj66, 22 September 2026: "If you code Opus 5.5 will be outstanding." and "If you are doing knowledge work, research or just regular office stuff Astra and the new Sol Models might still be the better deal for you." LocoMod, 23 September 2026, opened with "Astra is an incredible model." before praising Opus 5.5's speed. Majinvegito123, 23 September 2026, called Opus 5.5 an "incremental improvement for huge cost savings" and not the "Astra killer" some had hoped for.
Claude Opus 5.5 vs GPT-6 Astra safety
Neither lab compares the two models on safety in what we read, and no outside evaluator we read has done so, so there is no ranking to give. What each lab reports about its own model differs in kind: Anthropic ships Opus 5.5 with classifiers that block some cyber, biology and frontier-AI requests and can hand them to older models, and OpenAI classes Astra as its first model at the Critical level of cybersecurity capability.
For Claude Opus 5.5, blocked cyber requests fall back to Claude Opus 4.8, and blocked biology and frontier-LLM requests to Claude Opus 5. That is automatic in Anthropic's apps; on the API, per the system card, "the developer must opt in to automatic fallbacks" (p. 48). Some Opus 5.5 benchmark scores above therefore include answers another model finished. The card also reports a regression on instructions planted in text a user pastes into their own prompt, about 2% of attempts at default effort and about 7.4% at max (p. 126). TuxSH, 22 September 2026, found Opus 5.5 often made useless by its cyber guardrails and wrote that they are worse than Astra's. The detail is in our Opus 5.5 review.
For GPT-6 Astra, OpenAI's safety overview states: "Astra is our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework." The same document says "GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol." We read that overview on 3 September and have read two sections of Astra's system card. Who OpenAI decided may use that capability is the subject of our analysis of the Critical cyber classification, and our Astra review has the rest. Our own writing workflow runs on Anthropic models, and the disclosure at the foot of this page says what that means for the comparison above.
What has changed since GPT-6 Astra launched?
GPT-6 Astra's prices have not moved; its outside scores have. OpenAI's model page on 23 September 2026 prints the same $10 and $50, $1.00 cached input, $12.50 cache writes and 272,000-token rule we read on 5 September.
- 3 September 2026: OpenAI releases GPT-6 Astra, and ARC Prize publishes its ARC-AGI-3 analysis of it the same day.
- 6 September 2026: our Astra review reads Artificial Analysis index v4.2 (Astra 55 at max), LMArena WebDev (Astra 1797, Fable 5.1 1762) and Epoch AI's model page (Capability Index 169).
- 22 September 2026: Anthropic releases Claude Opus 5.5 at $4 and $20, and Artificial Analysis and METR publish on it the same day. LMArena's Code board, dated that day, has Astra first at 1793 and no Opus 5.5 entry.
- 23 September 2026: Artificial Analysis's v4.3.2 data puts Astra at 52.67 at max. Epoch AI's data file gives Astra 166.6; we have not established whether that revises the 169 on its model page or is a separate figure.
Still open on the day of publication: Opus 5.5 results on LMArena, Epoch AI, ARC-AGI-3 and the Coding Agent Index; Astra's default effort; OpenAI's own documents behind its Terminal-Bench figures; and whether Anthropic's detailed pricing page carries a long-context rule. We will re-read both rate cards and each evaluator at least monthly, sooner when any of them publishes on either model, and date every change on this page.
Claude Opus 5.5 vs GPT-6 Astra: independent evaluations
Six organisations outside the two labs were checked on 23 September 2026, and three publish figures for both models. Artificial Analysis, on Intelligence Index v4.3.2 with every Opus 5.5 entry run under Anthropic's default fallback, scores Opus 5.5 at 57.62, 55.99, 53.58, 51.24 and 42.31 at max, xhigh, high, medium and low, for $5.98, $3.46, $1.82, $1.34 and $0.55 per index task, and GPT-6 Astra at 52.67, 52.39, 50.92, 49.57 and 45.78 for $3.26, $2.31, $1.73, $1.54 and $0.82. On its own Terminal-Bench 4.0 run (mini-swe-agent harness, three repeats) both score 59.6%, Opus 5.5 at max and xhigh and Astra at xhigh. Zapier's AutomationBench leaderboard v1.0.6 has Astra at max on 41.4% ($1.73 per task) and Opus 5.5 at max on 40.0% ($1.28), the Opus 5.5 run made during Anthropic's early access. ARC Prize has both at 98.5% on ARC-AGI-1 at high effort, Opus 5.5 ahead on ARC-AGI-2 at high (93.3% against 92.1%) and Astra ahead at max (95.0% against 91.7%), and ARC-AGI-3 results for Astra only: 62.7% at max on the Standard harness and 99.9% at high on the Provider Adapter harness. LMArena's Code Arena WebDev board has gpt-6-astra-max first at 1793 and no Opus 5.5 entry. Epoch AI publishes a Capability Index of 166.6 for Astra and no score for Opus 5.5, and METR published a qualitative Opus 5.5 summary with no numbers and no Astra comparison.
What we read
- Anthropic, Introducing Claude Opus 5.5, the announcement page: the benchmark table with its three numbered footnotes, the chart captions carrying the cost-per-task claims against GPT-6 Astra, and the price tableThe lab's own document · read Sep 23, 2026
- Anthropic, Claude Opus 5.5 System Card: the capability table and section 8 pages 174 to 180 and 209 to 212, the safeguards and fallback pages 12, 13 and 48, the pasted-text result on page 126 and the self-preference result on page 127The lab's own document · read Sep 23, 2026
- Anthropic, models overview in the Claude Platform docs, the Claude Opus 5.5 row: context window, maximum output, default effort, thinking and knowledge cutoffThe lab's own document · read Sep 23, 2026
- Anthropic, claude.com pricing page, the API table for Opus 5.5 and the batch, fast mode, US-only inference and cache TTL lines, served in US dollars from our locationThe lab's own document · read Sep 23, 2026
- OpenAI, model page for gpt-6-astra: list price, cached input, cache writes, the 272K rule, batch, flex and fast mode multipliers, context window, maximum output, knowledge cutoff and effort settings. Re-read in a browser for this pageThe lab's own document · read Sep 23, 2026
- OpenAI API pricing page, the Standard table with the long-context columns and the data-residency noteThe lab's own document · read Sep 5, 2026
- OpenAI, Safety overview: GPT-6 Astra, for the Critical cybersecurity classification and the monitorability sentenceThe lab's own document · read Sep 3, 2026
- Artificial Analysis, Claude Opus 5.5 launch article, dated 22 September 2026, including its Terminal-Bench 4.0 line on GPT-6 Astra and its output-token counts per index task for both modelsIndependent · read Sep 23, 2026
- Artificial Analysis, the five Claude Opus 5.5 model pages on Intelligence Index v4.3.2, one per effort setting, including the GPT-6 Astra comparator rows in the page dataIndependent · read Sep 23, 2026
- Artificial Analysis, intelligence-benchmarking methodology, for the Terminal-Bench 4.0 harness and the cost basisIndependent · read Sep 23, 2026
- Artificial Analysis, Coding Agent Index v1.5, with a GPT-6 Astra entry and no Opus 5.5 entryIndependent · read Sep 23, 2026
- Zapier, AutomationBench leaderboard version 1.0.6, the GPT-6 Astra and Claude Opus 5.5 rows and the domain leadersIndependent · read Sep 23, 2026
- ARC Prize, results page for Claude Opus 5.5, verified scores at five effort settings, with the GPT-6 Astra comparator values in the page dataIndependent · read Sep 23, 2026
- ARC Prize, leaderboard and blog, checked for an Opus 5.5 ARC-AGI-3 result and finding noneIndependent · read Sep 23, 2026
- LMArena (arena.ai), the Text, Code/WebDev, Vision, Document, Agent and Search boards and the changelog, checked for an Opus 5.5 entry and finding noneIndependent · read Sep 23, 2026
- Epoch AI, benchmarks hub and Capability Index data files, with a GPT-6 Astra score and no Opus 5.5 scoreIndependent · read Sep 23, 2026
- METR, its predeployment summary of Claude Opus 5.5 and its time-horizons pageIndependent · read Sep 23, 2026
- Reddit comments naming both Opus 5.5 and Astra, in r/ClaudeAI, r/codex, r/OpenAI, r/singularity, r/accelerate, r/opencode and r/artificial, read through a public archive searchIndependent · read Sep 23, 2026
- Hacker News comments naming both Opus 5.5 and Astra, read through the site's search APIIndependent · read Sep 23, 2026
What we did not read
- Most of OpenAI's documents on GPT-6 Astra, as of today. The only OpenAI page re-read for this comparison on 23 September 2026 is the gpt-6-astra model page. The API pricing page, including the data-residency note, was read on 5 September; the announcement and the safety overview on 3 September; the ChatGPT and business plan pages on 5 and 6 September. Any change to those pages since then is not reflected here.
- Almost all of the GPT-6 Astra system card. We have read sections 7 and 4.1.1 of its 26,988 words, on 5 September 2026, and nothing in this comparison rests on the rest.
- OpenAI's own reports of the two Astra figures Anthropic's table prints for Terminal-Bench 4.0 (57.9% at high) and Terminal-Bench-Science (64.6% at max). Both reach this page through Anthropic's announcement and system card, which attribute them to OpenAI; we have not checked them against an OpenAI document.
- Anthropic's detailed pricing page, linked from claude.com/pricing. The pricing page we read prints no long-context surcharge for Opus 5.5; we did not open the detailed page, so we do not say that no such rule exists. Anthropic's extended cache TTL pricing is unread too; every Opus 5.5 cache price here is for the 5-minute TTL.
- Any statement of GPT-6 Astra's default reasoning effort. The model page we re-read does not give one, and we found none elsewhere in what we read.
- The Opus 5.5 system card beyond the pages named in the source list. The full record of what we read of it is in our Opus 5.5 review.
- Cognition's FrontierCode results, Proximal's FrontierSWE results and the Terminal-Bench 4.0 and Terminal-Bench-Science public leaderboards. Every figure from those operators on this page comes through Anthropic's system card.
- Artificial Analysis's per-evaluation pages for GPT-6 Astra. The Astra values on this page come from the comparator data on the Opus 5.5 pages, so we have no Astra AA-LCR, speed or time-to-first-token figure.
- Cloud marketplace rate cards for either model on Amazon Bedrock, Google Cloud Vertex AI or Microsoft Azure, and every consumer or team subscription plan. This page compares API rates only.
- The AssBench comparison one Reddit user cited, and any leaderboard not named in the source list.
- Reddit and Hacker News comments beyond what the archive search and the search API returned on 23 September 2026. G2, Trustpilot and Capterra carry no reviews of model APIs, and Opus 5.5 was one day old.
- Both models. We ran no prompt through Claude Opus 5.5 or GPT-6 Astra for this page and report no result of our own.
- Anything published after 23 September 2026.
The governance and safety questions this model raises are argued separately, with the evidence, in Is GPT-6 Astra AGI? The Critical classification, and who decided who gets it. This page does not repeat that argument.
Disclosure. Our writing workflow runs on Anthropic models, including Claude Opus 5.5, one of the two models compared. The Opus 5.5 system card measures a small but statistically significant bias toward itself when reminded that it is Claude (0.07 points out of 10, p. 127). So every comparison on this page rests on published figures from the two labs and from outside evaluators, each named with its setting.
We run no hands-on tests. This comparison is built from the lab’s own published documents, independent evaluations by other organisations, and dated user reports, each named above with the date we read it. How we investigate →