Frontier model comparison · Anthropic and OpenAI · two rate cards, both launch posts, OpenAI's Sol appendix, ten outside evaluators and 19 dated user reports
Claude Opus 5.5 vs GPT-6 Sol: Sol is half the price per token, and the fight is at the effort setting
GPT-6 Sol lists at $2 and $10 per million tokens against Claude Opus 5.5's $4 and $20, and a cached read costs $0.20 on both. Per task the gap is wider than 2x at the same effort setting, and it shrinks or flips at a level score: Zapier has Sol at xhigh level with Opus 5.5 at high for 38% of the cost, while Cognition has Opus 5.5 at medium beating Sol at max for 39% of the cost. Both labs' documents, ten outside evaluators and 19 dated user reports, read on 23 and 24 September 2026.
Our verdict
Run Claude Opus 5.5 for coding agents and for any job that needs a score GPT-6 Sol does not reach at any effort setting, preferably at medium effort, where Cognition's FrontierCode has it ahead of every Sol setting for $0.80 a rollout; run GPT-6 Sol at xhigh for high-volume work that a mid-range score already clears, such as business automations, where Zapier has it level with Opus 5.5 at high for 38% of the cost per task.
Claude Opus 5.5 vs GPT-6 Sol specs
- Released
- Both on 22 September 2026: Claude Opus 5.5 per Anthropic's announcement, GPT-6 Sol per OpenAI's
- API model ID
- claude-opus-5-5 (Anthropic's models overview) / gpt-6-sol (OpenAI's model page)
- Context window
- Claude Opus 5.5: 1M tokens. GPT-6 Sol: 1,050,000 tokens
- Maximum output
- Claude Opus 5.5: 128K tokens. GPT-6 Sol: 128,000 tokens
- Knowledge cutoff
- Claude Opus 5.5: June 2026, stated as its reliable knowledge cutoff. GPT-6 Sol: 20 April 2026
- Effort settings and default
- Claude Opus 5.5: low, medium, high, xhigh and max, default medium, thinking always on. GPT-6 Sol: none, low, medium, high, xhigh and max, default medium
- Availability
- Claude Opus 5.5: all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure, per the announcement. GPT-6 Sol: the API, ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, and not yet ChatGPT's Chat, per OpenAI's announcement
Claude Opus 5.5 vs GPT-6 Sol pricing
- Input, per 1M tokens
- Claude Opus 5.5 $4.00 / GPT-6 Sol $2.00. Sol is 50% of Opus 5.5 (our division)
- Output, per 1M tokens
- Claude Opus 5.5 $20.00 / GPT-6 Sol $10.00. Sol is 50% of Opus 5.5 (our division)
- Cache read, per 1M tokens
- Claude Opus 5.5 $0.20 (5-minute TTL) / GPT-6 Sol $0.20 cached input. Identical
- Cache write, per 1M tokens
- Claude Opus 5.5 $5.00 / GPT-6 Sol $2.50. Both are 1.25x the vendor's own input rate
- Long prompts
- Claude Opus 5.5: the pricing page we read prints no long-context surcharge. GPT-6 Sol: prompts over 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request, which is $4.00 input, $0.40 cached input and $15.00 output (our arithmetic)
- Batch
- Claude Opus 5.5: 50% saving, so $2.00 and $10.00 (our arithmetic). GPT-6 Sol: Batch and Flex at 50% of Standard, so $1.00 and $5.00 (our arithmetic)
- Fast mode
- Claude Opus 5.5: $8.00 and $40.00, 2x standard, stated as up to 2.5x speed. GPT-6 Sol: 2x the applicable rates, so $4.00 and $20.00 at short context (our arithmetic)
- Regional processing
- Claude Opus 5.5: US-only inference at 1.1x input and output. GPT-6 Sol: regional processing adds 10% where available, and EU data residency is available only with Standard processing
- Rate cards read
- Claude Opus 5.5: claude.com pricing page, 23 September 2026, re-read unchanged 24 September 2026. GPT-6 Sol: OpenAI's gpt-6-sol model page, 24 September 2026
Rate card read at the vendor’s own documentation on Sep 24, 2026.
GPT-6 Sol lists at $2 per million input tokens and $10 per million output, half of Claude Opus 5.5's $4 and $20, and a cached read costs $0.20 on both. The effort setting decides the rest. On Zapier's AutomationBench, Opus 5.5 at high scores 33.03% for $0.71 a task and Sol at xhigh 33.2% for $0.27. On Cognition's FrontierCode, Opus 5.5 at medium beats Sol at max, 54.6% against 49.3%, for $0.80 against $2.07.
Neither lab's launch material sets these two models against each other, so every head-to-head figure below comes from an outside evaluator that ran both, named with each model's effort setting. Two cautions apply throughout. Effort names belong to each vendor, so medium on one API does not buy the same amount of thinking as medium on the other. And several Opus 5.5 scores include tasks that an older Claude model finished after a safeguard refused them.
Neither model was run for this page. OpenAI's documents and the outside leaderboards were read on 24 September 2026, Anthropic's documents on 23 September with its pricing page re-read unchanged on 24 September, and the user comments below were posted between 22 and 24 September. Each model's full rate card and safety record sit in its own review; this page covers where the two differ.
Which is better, Claude Opus 5.5 or GPT-6 Sol?
Claude Opus 5.5 scores higher on every outside leaderboard we read that lists both, and GPT-6 Sol costs less per task on nearly every one that prints a cost. On Artificial Analysis's Intelligence Index v4.3.2, Opus 5.5 leads at each matched effort setting by 8 to 12 points, and Sol's best, 47.53 at max, sits below Opus 5.5 at its default medium (51.24). Sol's case is volume: where both models can reach a score, Sol usually reaches it for less.
Six leaderboards from five publishers carried both models on 24 September: Artificial Analysis's Intelligence Index and Coding Agent Index, Zapier's AutomationBench, Cognition's FrontierCode, UC Berkeley's Agents' Last Exam and LMArena's WebDev board. Opus 5.5 is ahead on all six at Sol's best setting. At the same named effort, Opus 5.5 costs 4.2 to 6.5 times as much per Artificial Analysis index task and 2.7 to 4.2 times as much per Zapier task, our division, both well past the 2x list gap. FrontierCode breaks the pattern: Opus 5.5 at medium costs $0.80 a rollout against Sol's $0.77 and scores 8.7 points higher.
So the choice turns on the score a job needs. If Sol at some setting clears your bar, it will usually clear it more cheaply. If the bar sits above what Sol reaches at max, only Opus 5.5 is in the running from these two, and GPT-6 Astra, OpenAI's $10 and $50 flagship, becomes the OpenAI model to weigh, as our Opus 5.5 and Astra comparison sets out. Where a single benchmark below matches your work, it outweighs this summary.
Why didn't OpenAI or Anthropic compare these two directly?
Both launched on 22 September 2026, and neither launch post names the other model. OpenAI compares GPT-6 Sol with Claude Opus 5, Fable 5 and Fable 5.1; Anthropic compares Claude Opus 5.5 with GPT-5.6 Sol and GPT-6 Astra. OpenAI's headline, "GPT-6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5's cost per task.", checks out on Zapier's page ($0.27 against $3.05), but Opus 5 is the model Opus 5.5 replaced that day. The same Zapier table has Opus 5.5 at max 9.27 points above Sol's best row.
Three of OpenAI's own Sol figures, 68.8% on DeepSWE, 60.5% on OSWorld 2.0 and 56.4% on Agents' Last Exam, did not appear on those benchmarks' leaderboards when we checked on 24 September, so they rest on OpenAI's runs alone. Key_Reading_9664 on r/codex, 22 September 2026, read the timing as a way to leave Opus 5.5 out, "so they didn't have to include it in benchmark comparisons". Neither lab says anything about the timing in what we read.
How does GPT-6 Luna change the choice?
Only for bulk work that a much lower score can handle. GPT-6 Luna lists at $0.10 and $0.50 per million tokens, 2.5% of Claude Opus 5.5's rates by our division, and costs $0.07 per Artificial Analysis index task at max for a score of 37.26, against Opus 5.5 at low on 42.31 for $0.55. On FrontierCode, Luna at max scores 42.4% for $0.10 a rollout and Opus 5.5 at low 47.3% for $0.40.
Two warnings come with the price. OpenAI's safety appendix says some of Luna's jailbreak results "may reflect a broader tendency to refuse requests, including legitimate ones", and pjankiewicz on Hacker News, 23 September 2026, reported on their own 36 agent scenarios that "the results were 30/36 for gpt 6 luna, and 33/36 for gpt 5.6 luna." Luna's full record is in our GPT-6 Sol review.
Is GPT-6 Sol cheaper than Claude Opus 5.5?
Per token, yes: GPT-6 Sol costs half of Claude Opus 5.5 on input, output and cache writes ($2, $10 and $2.50 per million tokens against $4, $20 and $5). Cached reads cost the same on both, $0.20 per million, and Anthropic says cache reads "make up the majority of agentic and coding work costs", so the more of a prompt that comes from cache, the further Sol's share of the bill rises above 50%. Rates are from OpenAI's model page, read 24 September 2026, and Anthropic's pricing page, read 23 September and re-read unchanged on 24 September; the Sol rate card in our review and the Opus 5.5 rate card in our review list every mode.
The 2x ratio holds across the modes both vendors publish. Batch halves both, to $1 and $5 against $2 and $10, and fast mode doubles both, which puts Sol in fast mode at $4 and $20, Opus 5.5's standard rate, our arithmetic from each vendor's multiplier. Regional rules differ in form and land at 10% each: Anthropic charges 1.1x for US-only inference, and OpenAI adds 10% for regional processing where available.
What happens to GPT-6 Sol pricing above 272K tokens?
Above 272,000 input tokens, GPT-6 Sol bills the whole request at twice its input and cache rates and 1.5 times its output rate: $4 input, $0.40 cached input and $15 output per million, our arithmetic. On a mostly cached prompt that erases Sol's per-token advantage over Claude Opus 5.5, because the Anthropic pricing page we read prints no long-context surcharge.
| Prompt size | Claude Opus 5.5 | GPT-6 Sol | Sol as a share of Opus 5.5 |
|---|---|---|---|
| 100,000 tokens | $0.296 | $0.156 | 53% |
| 271,000 tokens | $0.330 | $0.190 | 58% |
| 273,000 tokens | $0.331 | $0.331 | 100% |
| 500,000 tokens | $0.376 | $0.422 | 112% |
Our arithmetic from each vendor's list rates. Each request carries 20,000 new input tokens and 10,000 output tokens, with the rest of the prompt read from a cache written by an earlier call; cache writes and tool fees are left out. Opus 5.5 is priced with no long-context surcharge, which is what the page we read shows.
Crossing the line by 2,000 tokens raises Sol's bill on this request by 74% and Opus 5.5's by less than a tenth of a cent. The table gives both models the same 10,000 output tokens, which favours Opus 5.5, since at a matched effort setting Artificial Analysis counts three to four times as many output tokens per task from it. And the Claude pricing page links to a detailed pricing page we did not open, so an Opus 5.5 long-context rule may exist. For prompts that carry whole repositories, count input tokens before a call goes to Sol.
Claude Opus 5.5 vs GPT-6 Sol cost per task
At the same named effort, Claude Opus 5.5 costs 4.2 to 6.5 times as much as GPT-6 Sol per Artificial Analysis index task and scores 8 to 12 points higher. The multiple runs past the 2x list gap because Opus 5.5 writes three to four times as many output tokens per task at each setting.
| Effort setting | Opus 5.5 index score | Opus 5.5 cost per index task | Sol index score | Sol cost per index task |
|---|---|---|---|---|
| max | 57.62 | $5.98 | 47.53 | $1.06 |
| xhigh | 55.99 | $3.46 | 44.10 | $0.53 |
| high | 53.58 | $1.82 | 42.82 | $0.37 |
| medium | 51.24 | $1.34 | 39.78 | $0.25 |
| low | 42.31 | $0.55 | 33.90 | $0.13 |
Artificial Analysis Intelligence Index v4.3.2, read 24 September 2026, unrounded scores from the data behind the publisher's model pages. Every Opus 5.5 entry runs with Anthropic's default fallback; no Sol entry carries a fallback label. Sol also has a no-reasoning entry at 28.09 for $0.33, with no Opus 5.5 counterpart because Opus 5.5's thinking cannot be switched off.
Two cross-setting pairs say more than the matched rows. Sol at high (42.82 for $0.37) scores level with Opus 5.5 at low (42.31 for $0.55) at 67% of the cost. And Opus 5.5 at medium, its default (51.24 for $1.34), outscores Sol at max (47.53 for $1.06) for 26% more per task, with no Sol setting scoring higher. Our division throughout.
The publisher counts about 119,200 output tokens per index task for Opus 5.5 at max against 31,200 for Sol, and 25,700 against 6,500 at medium. Ormusn2o on r/singularity, 22 September 2026, wrote "According to AA, tokens per task for Opus 5.5 is about 5x of Sol 6.0"; the per-task counts put it at three to four times at a matched setting, our division.
Is GPT-6 Sol at max effort worth the extra cost?
On the outside numbers, rarely. Zapier has GPT-6 Sol at max on 32.0% for $0.34 a task, below Sol at xhigh on 33.2% for $0.27, and on Artificial Analysis's index Sol at max costs twice as much per task as at xhigh ($1.06 against $0.53) for 3.43 more points. _j on the OpenAI developer forum, 22 September 2026, read the launch charts the same way, describing "gpt-6-sol max running its reasoning up to max actually making no headway in the problems and outputting worse for the increased cost".
The same logic applies to Claude Opus 5.5 on coding: FrontierCode scores it highest at medium, 54.6%, against 54.4% at max for almost eight times the cost per rollout. Both labs default to medium.
OpenAI says effort changes on Sol now keep earlier context available for cache reuse, which MaitoSnoo on r/codex, 22 September 2026, passed on as "you can change the reasoning level for GPT-6 Sol and Luna mid-session without invalidating the cache". A Codex issue filed by 2c67cc18 on 24 September 2026 reports that after a mid-thread switch to high, Sol reasoned at least 1.26 times less than in a thread started at high, under an override flag the reporter describes as under development and off by default. Until that is settled, start a Sol thread at the effort you mean to pay for.
Claude Opus 5.5 vs GPT-6 Sol for coding
Claude Opus 5.5 leads GPT-6 Sol on every coding leaderboard that lists both, and on Cognition's FrontierCode it does so for almost the same money: 54.6% against 45.9% at medium, $0.80 against $0.77 per rollout. On Artificial Analysis's Coding Agent Index v1.5, run at max only, Opus 5.5 in Claude Code scores 66.0 and Sol in Codex 56.7, at $13.04 against $2.99 per task.
| Effort setting | Opus 5.5 score | Opus 5.5 cost per rollout | Sol score | Sol cost per rollout |
|---|---|---|---|---|
| max | 54.4% | $6.19 | 49.3% | $2.07 |
| xhigh | 51.4% | $2.25 | 48.4% | $1.32 |
| high | 54.0% | $1.09 | 47.7% | $1.04 |
| medium | 54.6% | $0.80 | 45.9% | $0.77 |
| low | 47.3% | $0.40 | 37.3% | $0.43 |
Cognition, FrontierCode 1.1 Main, read 24 September 2026. Cost is mean spend per rollout. Opus 5.5 ran in Claude Code and Sol in Codex, each lab's own agent harness, and runs flagged for unfair internet use score zero.
At low effort Opus 5.5 is both cheaper and 10 points higher, and Sol's expensive settings fare worst against it: Sol at max costs $2.07 a rollout, nearly twice Opus 5.5 at high ($1.09), for 4.7 fewer points. On the coding index, Sol's one lead is DeepSWE, 69.0% against 68.4%; Opus 5.5 is ahead on Terminal-Bench 4.0 (63.1% against 43.4%) and SWE-Atlas-QnA (66.4% against 57.5%). Of Opus 5.5's 909 attempts there, 78 were routed to a fallback model after a safeguard refusal; Sol's record shows none.
LMArena's WebDev board, built from people's votes on generated web apps, has Opus 5.5 first at 1818 and Sol fifth at 1686, both at max, and Agents' Last Exam, a professional-workflow benchmark rather than a coding one, gives Opus 5.5 at max a 38.2% pass rate against Sol's best of 32.2%, at xhigh.
Is GPT-6 Sol faster than Claude Opus 5.5 on coding tasks?
Yes. On Artificial Analysis's Coding Agent Index an average agent run took 1,338 seconds for GPT-6 Sol and 3,867 for Claude Opus 5.5, both at max. Sol also streams faster at every setting the publisher times for both: at medium, 114.3 output tokens a second and a first token in 1.96 seconds, against 79.4 tokens a second and 21.37 seconds for Opus 5.5.
The first-token time includes thinking, the publisher prints no Opus 5.5 speed at max, and these live figures, read once on 24 September, drift from day to day. DistanceSolar1449 on r/OpenAI, 23 September 2026, put Sol at a similar rate: "GPT-6-Sol is roughly 110 tokens/sec."
Claude Opus 5.5 vs GPT-6 Sol for business automation
On Zapier's AutomationBench, Claude Opus 5.5 at max scores 42.47% against GPT-6 Sol's best of 33.2%, at xhigh. At a level score Sol costs 38% as much: Sol at xhigh reaches 33.2% for $0.27 a task and Opus 5.5 at high 33.03% for $0.71. Zapier's current Opus 5.5 rows include tasks that Anthropic's fallback models finished after a refusal.
| Effort setting | Opus 5.5 score | Opus 5.5 cost per task | Sol score | Sol cost per task |
|---|---|---|---|---|
| max | 42.47% | $1.44 | 32.0% | $0.34 |
| xhigh | 35.77% | $0.89 | 33.2% | $0.27 |
| high | 33.03% | $0.71 | 31.2% | $0.24 |
| medium | 29.53% | $0.65 | 26.9% | $0.21 |
| low | 24.2% | $0.51 | 21.2% | $0.19 |
Zapier AutomationBench leaderboard 1.0.6, all 112 rows read 24 September 2026. Each task passes or fails on the final state of its environment, with no model grader. Every Opus 5.5 row carries the label "default fallbacks".
Which set of rows applies depends on how you call Claude Opus 5.5. Zapier re-scored it with fallback between 23 and 24 September, and on the earlier rows, which per Zapier's note predate the fallback reruns, Opus 5.5 at max read 40.0%, and at high 32.0%, which is below Sol at xhigh (33.2%). On the API, Anthropic's system card says "the developer must opt in to automatic fallbacks" (p. 48), so an integration that has not opted in is closer to the no-fallback rows, where Sol at xhigh ($0.27) edges Opus 5.5 at high ($0.65) for about 42% of the cost per task, our arithmetic.
Opus 5.5 leads Zapier's Sales and Operations domains, and no Sol row leads or comes second in any domain.
Claude Opus 5.5 vs GPT-6 Sol context window
GPT-6 Sol takes 1,050,000 tokens of context and Claude Opus 5.5 1M, a gap of about 5% that matters only for prompts near the limit, and both cap output at 128,000 tokens. The larger difference is price, since Sol reprices the whole request above 272,000 input tokens, as the table above shows.
On long-document reasoning the two are close: Artificial Analysis's AA-LCR has Opus 5.5 at 84.7% and Sol at 83.7% at max, and 84.3% against 82.3% at medium. Sol's knowledge cutoff is 20 April 2026, and Anthropic gives June 2026 as Opus 5.5's reliable cutoff. Sol can also run with no reasoning, which Opus 5.5 cannot. As an orchestrator, joseffb78 on the OpenAI developer forum, 23 September 2026, found "Sol High is decent" but wanted more context window for the role.
Is GPT-6 Sol safer than Claude Opus 5.5?
No one has compared them. Neither lab's material nor any outside evaluator we read sets the two side by side on safety, so there is no ranking to give. The labs also publish different amounts: Opus 5.5 has its own system card, while OpenAI's safety data for Sol is an appendix, added on 22 September 2026, to the GPT-6 Astra system card.
OpenAI treats GPT-6 Sol as High capability in cybersecurity and in biological and chemical risk, carries over the safeguards of GPT-5.6 Sol, and names no outside evaluator for it in the appendix. In one agentic test, among runs where the model found a message board of apparent peer agents, "GPT-6 Sol did so in 26% and 11% of samples, respectively": it tried to talk to them in 26% and carried out the specified unauthorised action in 11%, against 84% and 52% for GPT-5.6 Sol and no unauthorised action from GPT-6 Astra. The Sol review walks through the rest of that appendix.
Claude Opus 5.5 ships with classifiers that block some cyber, biology and frontier-AI requests and hand them to older models, Opus 4.8 for cyber and Opus 5 for the rest, automatically in Anthropic's apps and only by opt-in on the API. Its card reports a regression on instructions planted in text a user pastes into a prompt, acted on in about 2% of attempts at default effort and 7.4% at max (p. 126). Over-refusal is a live complaint: John-Lussier, in a Claude Code issue filed 22 September 2026, wrote "New default Opus 5.5 refuses to do any cyber work." despite membership of Anthropic's Cyber Verification Program. The full record is in our Opus 5.5 review. Our writing workflow runs on Anthropic models, and the disclosure at the foot of this page says what that means for the comparison.
What do users say about Claude Opus 5.5 vs GPT-6 Sol?
Most launch-week comments that name both models favour Claude Opus 5.5 on quality, two even on cost per task against the published figures, and some split the work between them. Across this page we link 19 posts from 19 accounts on Reddit, Hacker News, the OpenAI developer forum and GitHub, posted between 22 and 24 September 2026. They are self-selected posts from the first days, many about subscription plans, and they show where the argument sits without settling it.
Four compared the two on quality and chose Opus 5.5. Ok_Bite_67 on r/codex, 22 September 2026: "I used gpt 6 Sol earlier and it feels exactly the same as gpt 5.6 Sol. Meanwhile Opus 5.5 is a BEAST." Embarrassed_Adagio28 on r/NowInTech, 22 September 2026: "Opus 5.5 is significantly better than gpt 6 sol and even astra in most cases." qkamikaze on r/codex, 22 September 2026: "The results between Opus and Sol were so clear that i dropped my subscription within half an hour." ShadyShroomz on r/accelerate, 22 September 2026, found Opus "lower cost per task in my personal benchmark, and better performance" and concluded that "at the opus/sol tier opus is better in every way for my use cases."
ShadyShroomz's lower cost per task runs against the matched-setting figures from Artificial Analysis and Zapier, which put Opus 5.5 dearer at every setting; FrontierCode at low effort is the published case where it holds.
Two split the work by task. MrVirtuosoReality on r/singularity, 23 September 2026: "I'm basically using opus 5.5 to review and handle my most complex tasks at a better cost per task. More of the grunt impl work handed to gpt-6 sol (high) reasoning which is more than capable of speeding up my workflow with enough quality." ptj66 on r/singularity, 22 September 2026: "If you code Opus 5.5 will be outstanding." and "If you are doing knowledge work, research or just regular office stuff Astra and the new Sol Models might still be the better deal for you."
Part of the case against Sol is about its predecessor. tacomaster05 on r/ChatGPT, 22 September 2026, found Sol worse than GPT-5.6 on every task they gave it that isn't coding: "It is a massive downgrade and I'm switching back to Claude." Artificial Analysis's launch article says the new models' "Intelligence Index and Coding Agent Index scores remain level with GPT-5.6", and on its data Sol at max is 0.56 points above GPT-5.6 Sol at max.
One report favours Sol. matheusmoreira on Hacker News, 22 September 2026: "My benchmark also suggests even OpenAI's Sol models can match Fable performance at a fraction of the cost." Their comparison is with Anthropic's Fable, unversioned in the comment, and Anthropic says Opus 5.5 "performs at the level of Claude Fable 5.1 on most work", so on their workload the race may be closer than the leaderboards show.
Subscription reports mix plans with API prices. Da_ha3ker on r/codex, 22 September 2026, read the Codex pricing page as "66% the price for subs, not 50%", because Sol's message allowance rose by less than the list price fell. And haowang02, in a Codex issue filed 24 September 2026, reported that on one ChatGPT Pro account every model picked, Sol included, answered like GPT-5.6 Luna: "I'm getting one model under five different names". That is one account's report, and a reason to confirm which model served a session before judging it.
What has changed since Claude Opus 5.5 and GPT-6 Sol launched?
OpenAI's rate card was read on 24 September and Anthropic's on 23 and 24 September, with no price change between readings; three outside records have moved, all on the Opus 5.5 side.
- 22 September 2026: Anthropic releases Claude Opus 5.5 at $4 and $20, and OpenAI releases GPT-6 Sol at $2 and $10, 50% below GPT-5.6 Sol's promotional price, with a Sol and Luna appendix added to the GPT-6 Astra system card. Cognition adds both models to FrontierCode the same day.
- 23 September 2026: Zapier lists Opus 5.5 at max on 40.0%, with no fallback label. LMArena adds Sol to its WebDev board.
- 24 September 2026: Zapier's Opus 5.5 rows carry "default fallbacks" and read 42.47% at max, with medium and low rows added. Artificial Analysis's Coding Agent Index lists Opus 5.5 at 66.0, and LMArena's WebDev board, dated 23 September, has Opus 5.5 first with no changelog entry.
Still open: ARC Prize has no GPT-6 Sol result, DeepSWE and OSWorld 2.0 list neither model, METR has published nothing on Sol, Artificial Analysis runs its coding index at max only, and Agents' Last Exam prints no usable Sol cost. Both rate cards and every evaluator above get a fresh read at least once a month, and sooner if either model moves on any of them; each change is dated here.
Claude Opus 5.5 vs GPT-6 Sol: independent evaluations
Ten organisations outside the two labs were checked on 24 September 2026, and five of them publish figures for both models on six leaderboards. Artificial Analysis, Intelligence Index v4.3.2, scores Claude Opus 5.5 at 57.62, 55.99, 53.58, 51.24 and 42.31 at max, xhigh, high, medium and low for $5.98, $3.46, $1.82, $1.34 and $0.55 per index task, every entry run with Anthropic's default fallback, and GPT-6 Sol at 47.53, 44.10, 42.82, 39.78 and 33.90 for $1.06, $0.53, $0.37, $0.25 and $0.13. Its Coding Agent Index v1.5, run at max only, has Opus 5.5 in Claude Code at 66.0 ($13.04 per task, 78 of 909 attempts routed to a fallback model) and Sol in Codex at 56.7 ($2.99). Zapier's AutomationBench 1.0.6, with every Opus 5.5 row labelled default fallbacks, has Opus 5.5 at 42.47% at max ($1.44) and 33.03% at high ($0.71), and Sol at 33.2% at xhigh ($0.27) and 32.0% at max ($0.34). Cognition's FrontierCode 1.1 has Opus 5.5 at medium on 54.6% ($0.80 per rollout) and Sol at max on 49.3% ($2.07). UC Berkeley's Agents' Last Exam has Opus 5.5 at max on a 38.2% pass rate and Sol's best at 32.2% (xhigh), with no usable Sol cost. LMArena's WebDev board has Opus 5.5 first at 1818 and Sol fifth at 1686, both at max. Epoch AI's data files hold Furniture Assembly rows for both (83.3% and 58.3% at max) that Epoch has not written about, and no Capability Index score for either. ARC Prize lists no Sol result, DeepSWE and OSWorld 2.0 list neither model, and METR has published nothing on Sol.
What we read
- OpenAI, Introducing GPT-6 Sol and Luna: the price table against GPT-5.6 promotional pricing, the AutomationBench, Agents' Last Exam, FrontierCode, DeepSWE and OSWorld 2.0 claims with their named Claude comparators, and the availability and evaluation footnotes. Read in a browserThe lab's own document · read Sep 24, 2026
- OpenAI, models overview in the API docs, the GPT-6 Sol, Luna and Astra rowsThe lab's own document · read Sep 24, 2026
- OpenAI, model page for gpt-6-sol: effort settings and default, context window, maximum output, knowledge cutoff, list and cached prices, cache writes, the 272K rule, regional processing, batch, flex and fast mode. Read in a browserThe lab's own document · read Sep 24, 2026
- OpenAI, model page for gpt-6-luna, for the Luna prices and multipliersThe lab's own document · read Sep 24, 2026
- OpenAI, GPT-6 Astra System Card: the appendix on GPT-6 Sol and GPT-6 Luna added on 22 September 2026 (section 11 in the web version, A in the PDF) and the change logThe lab's own document · read Sep 24, 2026
- Anthropic, Introducing Claude Opus 5.5, the announcement page: the benchmark table and its footnotes, including the AutomationBench fallback note, and the price tableThe lab's own document · read Sep 23, 2026
- Anthropic, Claude Opus 5.5 System Card: the safeguards and fallback pages 12, 13 and 48, the pasted-text result on page 126 and the self-preference result on page 127The lab's own document · read Sep 23, 2026
- Anthropic, models overview in the Claude Platform docs, the Claude Opus 5.5 row: context window, maximum output, default effort, thinking and knowledge cutoffThe lab's own document · read Sep 23, 2026
- Anthropic, claude.com pricing page, the API table for Opus 5.5 and the batch, fast mode, US-only inference and cache TTL lines, served in US dollars from our location, read on 23 September and re-read unchanged on 24 September 2026The lab's own document · read Sep 24, 2026
- Artificial Analysis, GPT-6 Sol and Luna launch article, dated 22 September 2026Independent · read Sep 24, 2026
- Artificial Analysis, the six GPT-6 Sol model pages on Intelligence Index v4.3.2, one per effort setting plus the non-reasoning entry, with scores, cost per index task, output tokens, speed and time to first tokenIndependent · read Sep 24, 2026
- Artificial Analysis, the five Claude Opus 5.5 model pages on Intelligence Index v4.3.2, re-read on the same day as the Sol pagesIndependent · read Sep 24, 2026
- Artificial Analysis, intelligence-benchmarking methodology, for the index version, harnesses and judgesIndependent · read Sep 24, 2026
- Artificial Analysis, Coding Agent Index v1.5, with entries for Claude Opus 5.5 in Claude Code and GPT-6 Sol in Codex, both at maxIndependent · read Sep 24, 2026
- Zapier, AutomationBench leaderboard 1.0.6, all 112 rows including every Claude Opus 5.5 and GPT-6 Sol effort setting, the fallback note and the domain leaders. The Opus 5.5 rows were also read on 23 September 2026, before they were re-scoredIndependent · read Sep 24, 2026
- Cognition, FrontierCode 1.1, all reasoning levels for Claude Opus 5.5 and GPT-6 Sol, with cost per rollout and harness, and the changelog. Read in a browserIndependent · read Sep 24, 2026
- UC Berkeley RDI, Agents' Last Exam leaderboard, version ALE-V1, every effort for GPT-6 Sol and the Claude Opus 5.5 row. Read in a browserIndependent · read Sep 24, 2026
- LMArena (arena.ai), the Code Arena WebDev board and the changelog, with the other boards checked for both modelsIndependent · read Sep 24, 2026
- Epoch AI, benchmarks hub and data files, with no Capability Index score for either modelIndependent · read Sep 24, 2026
- ARC Prize, results pages, checked for GPT-6 Sol (no page) and Claude Opus 5.5Independent · read Sep 24, 2026
- Datacurve, DeepSWE leaderboard, checked for both models and listing neitherIndependent · read Sep 24, 2026
- XLang, OSWorld 2.0 leaderboard results file, checked for both models and listing neitherIndependent · read Sep 24, 2026
- METR, blog and time-horizons page, checked for GPT-6 Sol and finding nothingIndependent · read Sep 24, 2026
- Reddit comments on GPT-6 Sol and Claude Opus 5.5 in r/codex, r/OpenAI, r/singularity, r/ChatGPT, r/NowInTech, r/accelerate and r/ClaudeAI, read through a public archive searchIndependent · read Sep 24, 2026
- Hacker News comments on GPT-6 Sol and Luna, read through the site's search APIIndependent · read Sep 24, 2026
- OpenAI developer community thread announcing GPT-6 Sol and GPT-6 Luna, read through the forum's own data feedIndependent · read Sep 24, 2026
- GitHub issues filed against openai/codex and anthropics/claude-code in the launch weekIndependent · read Sep 24, 2026
What we did not read
- Anthropic's announcement, system card and models overview on 24 September 2026. They were read on 23 September (the pricing page was re-read unchanged on 24 September); any other change since then is not reflected here.
- Anthropic's detailed pricing page, linked from claude.com/pricing. The pricing page we read prints no long-context surcharge for Opus 5.5; we did not open the detailed page, so we do not say that no such rule exists. Anthropic's extended cache TTL pricing is unread too; every Opus 5.5 cache price here is for the 5-minute TTL.
- The GPT-6 Astra system card outside the Sol and Luna appendix and the change log, and the GPT-5.6 system card whose safeguards the appendix adopts by reference.
- The Claude Opus 5.5 system card beyond the pages named in the source list. The full record of what we read of it is in our Opus 5.5 review.
- OpenAI's post on prompt caching for GPT-6, the ChatGPT plan pages and the Codex pricing page that users cite for message allowances.
- OpenAI's own runs behind its Sol figures on DeepSWE (68.8%), OSWorld 2.0 (60.5%) and Agents' Last Exam (56.4%), none of which appears on the benchmark owners' pages we read.
- Press coverage of the two launches, including TechRepublic's comparison, which blocked automated access.
- Cloud marketplace rate cards for either model on Amazon Bedrock, Google Cloud Vertex AI or Microsoft Azure, and every consumer or team subscription plan. This page compares API rates only.
- Reddit, Hacker News, forum and GitHub posts beyond what the archive search and the search APIs returned between 22 and 24 September 2026. G2, Trustpilot and Capterra carry no reviews of model APIs, and both models were two days old.
- Both models. We ran no prompt through Claude Opus 5.5 or GPT-6 Sol for this page and report no result of our own.
- Anything published after 24 September 2026.
Disclosure. Our writing workflow runs on Anthropic models, including Claude Opus 5.5, one of the two models compared. The Opus 5.5 system card measures a small but statistically significant bias toward itself when reminded that it is Claude (0.07 points out of 10, p. 127). So every comparison here rests on published figures, each named with its publisher and setting.
We run no hands-on tests. This comparison is built from the lab’s own published documents, independent evaluations by other organisations, and dated user reports, each named above with the date we read it. How we investigate →