AI Tools Police
Reader-supported — we may earn a commission from links, at no cost to you. Rankings are never sold. How we investigate →

Frontier model review · OpenAI · the rate card, the rate-limit table, both plan pages, four outside evaluations and 19 dated user reports

GPT-6 Astra review: the price and the score both depend on a setting nobody prints

OpenAI publishes one headline rate for GPT-6 Astra and six different prices for a million output tokens. Four outside evaluators published within three days of launch, and every number each of them reports belongs to a reasoning tier, a harness or an index version. The rate card, the full rate-limit ladder, twelve published ARC-AGI-3 results, both plan pages line by line and 19 dated user reports, read between 3 and 6 September 2026.

By Mucahit Kaya · Founder and EditorSep 6, 2026~18 min readgpt-6-astra

Our verdict

GPT-6 Astra ranks first on Epoch AI's capability index, first in LMArena's WebDev arena and second on Artificial Analysis v4.2 at max effort, at $10 and $50 per million tokens. GPT-5.6 Sol's $4 and $20 is promotional to 21 November 2026. Every headline number on this model moves with a setting, so price the workload you would actually move, at the tier you would actually run it.

Published specification

Context window
1,050,000 tokens
Maximum output
128,000 tokens
Knowledge cutoff, as printed on the model page
30 April 2026, read 5 September 2026
Reasoning effort settings
low, medium, high, xhigh, max. Every published score for this model belongs to one of them, or to a run with reasoning off.
Modalities
Text in and out. Image input. Audio and video listed as not supported.
Fine-tuning
Not supported
Tools via the Responses API
Web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, tool search
What OpenAI says it is for
"[B]uilt for the hardest end-to-end work. Use it for complex reasoning, coding, computer use, research, and document creation."

What it costs to run

Standard, short context
$10.00 per 1M input tokens, $50.00 per 1M output, $1.00 cached input
Standard, long context
$20.00 per 1M input tokens, $75.00 per 1M output, $2.00 cached input. Long context begins above 272,000 input tokens.
The 272,000-token rule
"Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request."
Cache writes
$12.50 per 1M at short context and $25.00 at long context, the stated 1.25x of the uncached input rate in each case
Batch, Flex and Fast mode
"Batch and Flex are priced at 50% of Standard rates. Fast mode is priced at 2x the applicable rates."
The six-way spread
One million output tokens costs $25.00 to $150.00 on this one model depending on mode and prompt length. Our arithmetic from OpenAI's published rates and its own stated multipliers.
Data residency endpoints
10% uplift for eligible models released on or after 5 March 2026
Amazon Bedrock
"OpenAI models in Amazon Bedrock are billed through AWS and may differ from direct OpenAI pricing."
API rate limits, as published
Free tier listed as not supported. Tier 1: 500 requests and 500,000 tokens per minute. Tier 2: 5,000 and 1,000,000. Tier 3: 5,000 and 2,000,000. Tier 4: 10,000 and 4,000,000. Tier 5: 15,000 and 40,000,000.
ChatGPT consumer plans
Free $0, Go $6, Plus $20, Pro from $120 per month, read on ChatGPT's own pricing page 6 September 2026. A subscription buys models inside an app and not the rates above.
ChatGPT Business and Enterprise seats
Standard seat $20 per month billed annually or $25 billed monthly; premium seat $100 annually or $125 monthly; Business sold for teams of 2 to 200; Enterprise quoted on contact, with credit-based and token-based options both mentioned. Read 5 September 2026 in the Turkish rendering that page serves us.
The comparator, and its expiry
GPT-5.6 Sol at $4.00/$20.00 is promotional pricing, "available at least through November 21, 2026"

Rate card read at the vendor’s own documentation on Sep 5, 2026.

GPT-6 Astra lists at $10 per million input tokens and $50 per million output at short context, and that pair is the cheapest of six prices OpenAI publishes for the same model. Cross 272,000 input tokens and the whole request reprices upward. Send the same job in Batch or Flex and it halves. Four organisations outside OpenAI published measurable results within three days of launch, and each of them reports a figure that moves with a reasoning setting, a harness or an index version. This review is about which setting each number belongs to, because that is the part that decides the invoice.

A word on how it was put together, because it bounds what the page can claim. We have not run this model. Everything here is read from documents, each carrying the date we opened it: OpenAI's announcement and its two safety pages on 3 September 2026, read by hand in a browser because that host refuses automated retrieval; the API rate card, the model page and the business pricing page on 5 September; two sections of the system card, out of 26,988 words, also on 5 September; ChatGPT's consumer pricing page, the four independent evaluations and sixteen Hacker News threads on 6 September. The full register is at the foot of this page, printed next to the longer list of what we did not open.

What the 2.5x actually compares

GPT-6 Astra's published rate is two and a half times GPT-5.6 Sol's, and that comparison has an expiry date printed under the same table.

The price side first, in full, so it can be cited briefly afterwards. On OpenAI's rate card, read 5 September 2026, gpt-6-astra is $10.00 per million input tokens and $50.00 per million output at short context. gpt-5.6-sol is $4.00 and $20.00. That is 2.5x on both sides, our division of two published rates. The qualifier sits directly beneath: "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026." Sol's $4 and $20 is a promotion with a stated end date, not a standing price, so the multiple is a fact about this quarter.

ModelInput, short contextOutput, short contextInput, long contextOutput, long context
gpt-6-astra$10.00$50.00$20.00$75.00
gpt-5.6-sol$4.00$20.00$8.00$30.00
gpt-5.6-terra$2.00$12.00$4.00$18.00
gpt-5.6-luna$0.20$1.20$0.40$1.80

OpenAI's published Standard-tier rates per one million tokens, read 5 September 2026. The Sol row is the promotional rate described above.

Which Sol price belongs in the comparison was itself argued out in public during launch week. simonw, on 5 September 2026, settled it by pointing at OpenAI's own changelog: $5 and $30 is the list price and $4 and $20 is the promotion that replaced it. We have not opened that changelog. Measured against a $5 and $30 list price, Astra is 2x on input and 1.67x on output, our calculation, and no better than the changelog figure underneath it. A budget built on the 2.5x and a budget built on the 2x are different budgets, and one of them changes on 21 November 2026.

What the extra money buys is the part the outside evaluators do not agree on, and the disagreement is only legible with the version attached. Artificial Analysis publishes an Intelligence Index and revised it between our two readings. On v4.1.1, in its launch article of 3 September 2026, Astra scored 61 and so did GPT-5.6 Sol. On v4.2, read on the model pages on 6 September 2026, Astra at max effort scores 55, Sol at max scores 51 and ranks fourteenth of 202. Those are two index versions, not one score that moved.

There is a second layer under that, and it is the one most commonly flattened. Artificial Analysis does not publish one number for this model on v4.2. It publishes six: 55 at max effort, 54 at xhigh, 53 at high, 52 at medium, 49 at low, and 48 for a run with reasoning off. The spread from top to bottom is seven index points, our subtraction, on one model, one index version, one publisher, one day. Anyone quoting "the Artificial Analysis score" for GPT-6 Astra is quoting one of six numbers without saying which.

Epoch AI's Capability Index puts Astra at 169, with a 90% confidence interval of 165 to 174 and a rank of first out of 267 models, dated 3 September 2026 and read on the 6th. Two qualifiers on it are Epoch's own, and both matter more than the rank. The headline is published as "Best score across settings", an aggregate over seven settings and not a single configuration. And Epoch writes that the figure "is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend". A record that lands on an existing trend line is a different claim from a break in it.

For a buyer the instruction narrows to one line. If the work you would move is work GPT-5.6 Sol already finishes, nothing in the four evaluations below argues for paying the premium to finish it again. If it is work Sol fails, the case has to be built on your own traffic, because none of these organisations measured your traffic.

One observation from the same rate card, stated once and not built on: gpt-6-astra does not appear in the Cyber models table on that page, whose note says the two Daybreak aliases "currently point to gpt-5.6-sol and gpt-5.6-cyber, respectively", read 5 September 2026. The access programme itself, and the question of who OpenAI decided may hold a capability it classified Critical, is argued with the documents in our companion piece and is not re-run here.

Where Claude Fable 5.1's rate card would go

Claude Fable 5.1 is the model most readers arrive wanting priced against this one, and this page does not price it. We did not open Anthropic's pricing page at any point in this reading, so there is no dated per-token figure for Fable 5.1 in our register and no Astra-to-Fable price multiple of ours anywhere below. An estimate would look like the missing row and would not be one.

What exists instead is one reader's arithmetic, carrying their own caveat. slopinthebag, 5 September 2026 wrote that "astra high is also 3x cheaper than opus max at basically the same intelligence", adding that "benchmarks are really fuzzy with llms". That multiple stays theirs. What we can compare is capability, because a third party published both sides on one index: on Artificial Analysis v4.2, read 6 September 2026, Claude Fable 5.1 is first of 202 at 57 and Astra at max effort is second at 55. On the earlier v4.1.1, Fable was 66 and Astra 61. Astra trails Fable on both versions, by five points on the older one and by two on the newer, each gap read inside its own version. That is the shape of the gap without a price attached to it, which is as far as our reading goes.

Where Gemini 3.8 Flash would go

Gemini 3.8 Flash is the third comparison people ask for, and this page cannot supply it either, for a reason worth printing rather than skipping. We opened no Google pricing page and no Google benchmark page, so no Gemini price and no Gemini score of ours appears here.

There is a wrinkle behind that absence that a reader can check for themselves. OpenAI's own announcement page carries a benchmark comparison table, read 3 September 2026, and Gemini 3.8 Flash is named among its columns. Our record of that table holds five values per row against those six named columns, so we cannot say which value belongs to Gemini, and we have not re-read the table to resolve it. Publishing a number under the wrong column heading is the specific way this figure would go wrong, so nothing is published under it. That gap is not ours alone: the GPT-6 Astra write-up by Josphat Mutai at ComputingForGeeks, updated 5 September 2026 and read by us on the 6th, reports its own API measurements of caching and latency, and its benchmark comparison sets Astra against GPT-5.6 Sol, Claude Opus 5 and Claude Fable 5.1 with no Gemini row in it.

One score on this model is never one score

A benchmark result for GPT-6 Astra is a claim about a configuration, and ARC-AGI-3 is where that stops being a technicality and starts being an amount of money.

ARC Prize Foundation, which owns the benchmark, published its own evaluation on 3 September 2026 and we read it on the 6th. On its ARC-AGI-3 Semi-Private set it publishes twelve results for this model, six reasoning settings across two harnesses, each with a cost per task.

Reasoning settingARC Prize Standard harnessProvider Adapter harness
max62.7%, $26,09898.6%, $17,332
xhigh59.3%, $37,31798.4%, $18,147
high54.8%, $40,70599.9%, $18,817
medium38.6%, $48,09098.4%, $19,285
low17.5%, $38,16698.0%, $21,298
reasoning off35.2%, $49,79196.7%, $23,457

ARC Prize Foundation's published results and per-task costs on the ARC-AGI-3 Semi-Private set, published 3 September 2026, read 6 September 2026.

A word first on what those percentages are, because 99.9% invites the reading that the model solved 99.9% of the tasks. ARC Prize's own framing of the benchmark is that ARC-AGI-3 "challenges AI agents to adapt on the fly to novel interactive environments", and people are scored on it as well as models: the caption on OpenAI's chart, read 3 September 2026, puts the average human tester at 48%. ARC Prize's own statement of how a run's percentage is computed against that human baseline is not something we read, so this page treats 99.9% as a score on ARC Prize's scale and not as a count of tasks solved, and the gap is in the register at the foot of the page.

Three readings come out of that table, and none of them is the headline. The first is that the 99.9% OpenAI quotes belongs to high effort, not to max: the Provider Adapter figure at max is 98.6%. The second is that the score curve on the Standard harness is not a ladder. Low effort scores 17.5%, which is 17.7 points below the same harness with reasoning switched off entirely, our subtraction, and ARC Prize does not comment on it. The third is the one with a bill attached: cost runs backwards. On the Standard harness the run with reasoning off cost $49,791 per task and the run at max effort cost $26,098, a fall of 47.6%, our calculation. The same direction holds on the Provider Adapter column. On this benchmark the most expensive setting per token produced the cheapest result per task, because the model finished in fewer moves.

That is a cost finding, and it is the one worth carrying into a routing decision: the reasoning setting is not a price dial that only goes one way. Whether the Provider Adapter harness or the Standard harness is the fair one to quote is a different argument, and the same harness question argued against the launch documents is where we make it, not here.

The same discipline bites twice more on this page. Artificial Analysis reports 61 and 55 for one model because an index version moved between two readings, and LMArena's WebDev row states no reasoning effort for Astra while the entry directly beneath it carries max inside its own name. A score with no configuration is not a smaller claim than a score with one. It is an incomplete claim, incomplete in the direction of whoever is quoting it. That is the per-vendor version of a problem we have written about before, in what happens when the measuring stops at one vendor's own traffic.

The 272,000-token line reprices everything behind it

One million output tokens on this one model costs anywhere from $25.00 to $150.00, depending on the mode it is sent in and how long the prompt is. The figure a reader meets first is $50.00, which is the third-cheapest of the six.

The model page states the rules, verbatim, read 5 September 2026:

"Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request."

"Batch and Flex are priced at 50% of Standard rates. Fast mode is priced at 2x the applicable rates."

The phrase that costs money is for the full request. Crossing 272,000 input tokens does not price the excess higher. It reprices everything in that call: input, cached input, cache writes and output. A 271,000-token prompt and a 273,000-token prompt are not one percent apart on the invoice.

One million output tokensShort contextLong context, over 272K input
Batch or Flex, at 50%$25.00$37.50
Standard$50.00$75.00
Fast mode, at 2x$100.00$150.00

Our arithmetic, computed from OpenAI's published rates and its own stated multipliers, read 5 September 2026. Output tokens only, before input, caching or tool-call charges.

Every multiplier reconciles against the printed table, which is the part OpenAI does well. Long-context input at $20.00 is the stated 2x of $10.00, long-context output at $75.00 is the stated 1.5x of $50.00, and cache writes are the stated 1.25x of the input rate at both context lengths, $12.50 and $25.00. Nothing is concealed. The problem is that six published prices for one advertised model arrive under one headline figure, and the headline is the small one. We have made the general version of this argument in why a rate you cannot convert into finished work is not yet a price; the per-token version is not a missing number, it is six of them.

Where the threshold sits inside the advertised window is what gives it its edge. The model page gives a 1,050,000-token context window. 272,000 is 25.9% of it, our calculation. The cheaper rate covers roughly the first quarter of the context length the model is sold on, and the other three quarters sit past a line that reprices the whole call.

Caching is the lever that moves against all of this, and the published rates make the shape of it clear: cached input is $1.00 per million against $10.00 uncached, with a $12.50 write. ComputingForGeeks, updated 5 September 2026, reports measuring a 91.9% cost reduction on a repeated prompt through the API, alongside a measured ladder of latency, token count and cost across all five reasoning-effort levels. We did not reproduce that figure and we ran no calls of our own, so it is theirs and it is linked so the method can be read.

We looked for anyone reporting a bill from the 272,000-token rule and found nobody. Across sixteen Hacker News threads and roughly 2,900 comments read on 6 September 2026, searched for "272", "reprice", "surcharge", "1.5x output" and "2x input", there is no report of anyone crossing it. The nearest matches are 2025 comments about GPT-5's 272,000-token input limit, which is a different thing entirely. Three days after launch that is an absence of evidence, and it is the figure on this page we would most expect to change.

The practical move is to count input tokens before dispatch instead of after the invoice. Put a hard ceiling below 272,000 in whatever wraps the calls, or price the job at the long-context column from the start and enjoy being wrong.

Rate limits, from Tier 1 to Tier 5

The published rate-limit table for gpt-6-astra, read 5 September 2026, is the part of the access question that has real numbers attached, and it starts with a closed door.

Usage tierRequests per minuteTokens per minuteBatch queue limit
FreeNot supported
Tier 1500500,0001,500,000
Tier 25,0001,000,0003,000,000
Tier 35,0002,000,000100,000,000
Tier 410,0004,000,000200,000,000
Tier 515,00040,000,00015,000,000,000

OpenAI's published rate limits for this model, read 5 September 2026 on the model page.

Two things in that table are worth converting into working terms. Moving from Tier 1 to Tier 2 multiplies requests per minute by ten and tokens per minute by two, our division, so the first promotion buys concurrency far more than it buys throughput. Throughput is where the ladder ends up steep: Tier 5's token allowance is eighty times Tier 1's and its request allowance thirty times, again our division.

The token line binds first for long-context work, and it binds sooner than it looks. A single prompt at the 272,000-token threshold consumes 54.4% of one minute's token allowance at Tier 1, our calculation, so an account at the entry tier can send roughly one such request per minute and not two. The Tier 1 batch queue holds about five of them. That is the ceiling a first project actually meets, and it is a different constraint from the price per token that gets quoted.

The waitlist we could not find, and why Daybreak is not it

An API waitlist for GPT-6 Astra is the one thing in this section we cannot describe, and saying so is more useful than filling the space. We read the published tier limits above on the model page and found no OpenAI document explaining how an account moves from one usage tier to the next, no application form and no waitlist page. So this review states the limits and does not state the route, and if a document setting out that route exists, we have not read it.

Two things that are frequently merged should be held apart, because they are unrelated systems that both use the word access. The usage tiers above are the ordinary API ladder, keyed to an account's own usage and spend, and they apply to this model like any other. Daybreak Blue and Daybreak Red are something else: an approval programme for cyber-capability access, with its own eligibility and its own review. As of the rate card read 5 September 2026 the Daybreak aliases pointed at gpt-5.6-sol and gpt-5.6-cyber, not at this model. A reader told they need approval for one of these has not necessarily been told anything about the other, and the mechanics of the second belong to our companion analysis rather than to a rate-card page.

The hallucination figures OpenAI says are not hallucination rates

OpenAI's system card states, in its own document, that its factuality figures are not production hallucination rates, and any page quoting them as such is misreporting a number whose own source warns against it.

That document runs to 26,988 words. We read section 7 on hallucinations and section 4.1.1, on 5 September 2026, and none of the rest. Section 7 says:

"We evaluate factuality on de-identified ChatGPT conversations that users of our prior models have flagged as containing factual errors. These examples are intended to capture especially hallucination-prone cases, so their absolute error rates are expected to be much higher than the true error rates in production and should not be interpreted as hallucination rates observed in production."

Two metrics are reported there: whether the model makes any error at all, which OpenAI calls the response-level hallucination rate, and whether it reproduces the specific error the user flagged. On both, the document reports that "Astra makes substantially fewer factual errors than GPT-5.6 Sol and is significantly less likely to reproduce user-reported hallucinations", and that "These improvements are particularly pronounced at very low latency and reasoning settings."

The useful part of that is the direction and the setting. The improvement runs downward against the previous model on a sample built to be hard, and it is largest at low latency and low reasoning effort, which happens to be where a cost-conscious deployment lives. The trap is the absolute number, because the sample is a worst case by construction and the vendor says so first. Do not carry an absolute error percentage for this model into a risk register, whoever printed it. A figure that will mean anything for your domain has to be measured on your own traffic. We did not extract the values from that section's figure, because the caveat is the finding.

Section 4.1.1 carries a different kind of number, and one reading of it is ours. Its Table 1 scores safe completions on challenging prompts across eight categories for four model generations. Astra's column is the highest in seven of the eight, its lowest cell is the gore category at 0.898, and the one row where its margin over gpt-5.6-sol is close to nothing is the minors-related sexual-content row, 0.974 against 0.973, a margin of a tenth of a point where the seven other categories gain between 1.0 and 11.3 points. That is our reading of the vendor's own table and not a claim OpenAI makes. The vendor's framing of the same section runs the other way and is worth having: it reports "a Pareto improvement in the rate of handling harmful requests safely vs. helpfulness on legitimate requests" and states that Astra "is less likely to refuse harmless requests or add excessive or unnecessarily judgmental caveats", which for anyone who has fought a refusal loop is the more practical sentence.

One boundary on all of the above. Other coverage of this launch prints specific prompt-injection resistance and misalignment percentages from this system card. Those sit in sections we did not open, so they do not appear here, and the alignment and monitorability material from the accompanying safety overview is read against the launch documents in our companion piece rather than summarised twice.

If you pay for ChatGPT and not for tokens

A ChatGPT subscription and the per-token API rate card set out earlier on this page are two different products, and the plan pages make the price legible while leaving the model's place in them unclear.

The naming first, because it sends people in circles. There is no product called ChatGPT 6. The model is GPT-6 Astra, its API identifier is gpt-6-astra, and ChatGPT is the application it appears inside. A subscription buys a set of models in an app, billed monthly. It does not buy the per-token rates above.

One bound belongs in front of everything that follows. chatgpt.com/pricing sets its locale from the connection, and from here it serves Turkish, with a request for the English rendering redirecting back. So everything below was read in the Turkish rendering, on 6 September 2026, where the four plan prices are printed in US dollars, which is why they can be quoted here as dollars. We did not open the English rendering and cannot say what it prints, and every plan description below is our translation of the Turkish.

Read 6 September 2026, the page lists Free at $0 a month, Go at $6, Plus at $20 and Pro from $120. GPT-6 Astra appears on that page in exactly one place, the Models row of the plan comparison table, and in no individual plan's feature list.

Free runs to eight lines, and the first two are worth setting down in the Turkish they are printed in, because the second qualifies the first: "GPT-5.6 Luna ile sınırsız metin sohbeti", which we read as unlimited text chat with GPT-5.6 Luna, then "Sınırlı mesaj ve dosya yükleme hakkı", a limited allowance of messages and file uploads. The remaining six lines, in our English, are limited and slower image generation, limited voice chat, limited deep research, limited memory and context, limited access to Codex, and limited access to ChatGPT Work on the desktop. That last line names the product "ChatGPT Çalışma" where the Plus list leaves "ChatGPT Work" in English, which is a small demonstration of why we are not going to guess at what the English rendering says. Go at $6 sits between Free and Plus, and it is the only tier whose description mentions advertising at all. That line reads, in Turkish, "Bu plan reklam içerebilir." We translate it as: this plan may contain ads.

The two paid tiers are written as additions to the tier below them. Plus at $20 heads its feature list with "Go planındaki tüm özellikler ve:", everything in the Go plan plus what follows, and nine lines follow: advanced reasoning models with GPT-5.6, a larger allowance of messages and file uploads, more complex and accurate image generation, more extensive deep research, more extensive memory and context, projects, scheduled tasks and custom GPTs, wider Codex usage, wider access to ChatGPT Work on desktop, web and mobile, and early access to new features. Pro from $120 heads its own list with "Plus'taki tüm özellikler ve:", the same construction against Plus, and eight lines follow: 5 or 20 times more usage, Pro reasoning with GPT-5.6 Sol Pro, maximum Codex tasks, unlimited and faster image generation, maximum deep research, maximum memory and context, wider projects, tasks and custom GPTs, and a research preview of new features. Every English word in those two lists is ours. The page carries them only in Turkish.

Across those seventeen paid-tier lines the only models named are GPT-5.6 and GPT-5.6 Sol Pro, both previous generation, so the tier a reader would pick for better reasoning names a model that is not the one on this page. The decision available from that page on 6 September 2026 is a $20 or $120 decision about message volume and Codex headroom, and the Models row of the comparison table is the only place the page commits to anything about Astra.

SurfaceWhat it showedRead
OpenAI announcement pageAstra "is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS"3 September 2026
chatgpt.com/pricingAstra named in the Models row of the plan comparison table, and in no plan's own feature list6 September 2026
openai.com/business/pricingVisible plan comparison table listing GPT-5.6 Sol, Sol Pro, Terra, Luna and GPT-5 Thinking Mini, with no Astra row. The string Astra is present in the page's markup5 and 6 September 2026
developers.openai.com, model pageFree tier listed as not supported; Tier 1 to Tier 5 limits published5 September 2026
developers.openai.com, Cyber models tablegpt-6-astra absent; Daybreak aliases pointing at gpt-5.6-sol and gpt-5.6-cyber5 September 2026
Amazon BedrockNamed on OpenAI's pricing page, "billed through AWS and may differ from direct OpenAI pricing"5 September 2026

What each surface showed on the date we read it. The business-page finding is about the visible comparison table, not about the whole page.

What the business page states beyond that table, read 5 September 2026: a Business standard seat is $20 per month billed annually or $25 billed monthly; a premium seat is $100 annually or $125 monthly; Business is sold for teams of 2 to 200; and Enterprise is quoted on contact, with credit-based and token-based options both mentioned. The locale bound stated above applies to this page too, and it is the second of the two that carry it: the business page serves Turkish from our location as well, the seat figures just given are the dollar amounts printed in that rendering, and we did not open the English one. So the seat descriptions are our reading and not the vendor's English. The premium seat's line reads "5 saatlik kullanım limiti olmadan, standart kullanıcı hakkına göre 5 kat daha fazla kullanım", which we read as five times the standard seat's usage, without a five-hour usage limit. The Turkish attaches that limit to no tier by name, so taking it for the standard seat's is our reading of a line printed inside the premium seat's own feature list. The table finding above rests on none of that: model identifiers are not translated, so a comparison table naming five models and no row for this one names the same five in any rendering.

Two numbers on the business comparison table need separating from the model's own, because they are the ones a seat-holder actually meets. Its context rows are GPT Instant at 54,000 tokens and GPT Reasoning at 256,000. Those are product surfaces, and neither is the 1,050,000-token window on the model page. A plan does not necessarily hand you the window the API advertises.

Here is where the cheap route stops, stated plainly because it is the thing most likely to catch a reader out. There is no free way to reach this model. The published rate-limit table lists the free API tier as not supported, so an experiment has to start at Tier 1 with a funded account, and a ChatGPT subscription does not convert into API credit at any tier. Two people paying for plans said what that felt like in the first three days. kbrannigan, 5 September 2026: "After 15 message I burned through my 5 hour limits." The only five-hour usage window in anything we read is the line printed inside the premium seat's feature list above, and that comment does not say which plan it is on. forrestthewoods, 5 September 2026 put $10 of API credit against a fantasy-auction task and reported that "It spend $3.50 and then said 'this action would cause you to go above your spending limit'", then bought a $100 Codex Max subscription that included Astra and finished the job. Those are two people's bills, dated, and they are not a plan table. What they show is the shape of the decision: the model is reachable cheaply enough to try and not cheaply enough to try casually, and the tier that solves it is the one above whichever tier a reader is currently on.

The four outside evaluations, in full

Four organisations outside OpenAI published something measurable within three days of this launch, which is fast, and no two of them measured the same thing. One carries a funding relationship with OpenAI on the benchmark family in question, and disclosed it itself. A fifth, whose silence matters most, has published nothing.

EvaluatorWhat it publishedThe configuration attachedThe caveat
ARC Prize Foundation99.9% on ARC-AGI-3 Semi-Private, at $18,817 per taskhigh effort, Provider Adapter harnessTheirs: "we are not claiming that it is AGI", and the benchmark "has a tightly bounded scope and format"
Artificial AnalysisIntelligence Index 55, second of 202max effort, index version v4.2, one of six published numbersThe index was v4.1.1 three days earlier and the versions are not comparable
Epoch AICapability Index 169, first of 267"Best score across settings", seven settings aggregatedTheirs: the jump is "within our uncertainty range for the reasoning-era ECI trend". OpenAI funded FrontierMath
LMArena1797 in the WebDev/Code arena, firstNone statedThe entry below it is named claude-fable-5.1-max and states its own
METRNothing on this modelNot applicableChecked 6 September 2026

The independent evaluations this page rests on, each with the configuration its score belongs to. All read 6 September 2026, except the ARC Prize leaderboard, read 3 September.

ARC Prize is worth reading past its headline on a point that has nothing to do with harnesses. On its code-sandbox setup it states that results "should be understood as the combined performance of the model and its tools", and, in its own words, "Our testing participants did not have a code interpreter, scratch pad, etc." Any comparison to a human baseline on this benchmark is a comparison between a model holding tools and people who were not holding any. ARC Prize also calls Astra "a noticeable step-function change in frontier model capabilities", and that sentence belongs on the page as much as the qualifications do.

Underneath Artificial Analysis's index sit its per-component figures, and this page has to stop short on them for the same reason it stops short on Gemini. Our record of those components, read 6 September 2026, holds seven movements: an AA-Omniscience hallucination rate recorded as 92% to 51%, Humanity's Last Exam up six points, AA-Briefcase up roughly 80 Elo, GDPval-AA v2 down roughly 80 Elo, and regressions of two to three points on tau-cubed-Banking, SciCode and AA-LCR. What that record does not hold is what any of them is measured against, whether the previous model, the previous index version or another effort tier, and it does not hold the reasoning setting either side of a movement ran at. Comparative figures with no comparand attached are the exact claim this page spends its length arguing against, so they sit here as an unresolved record, nothing of ours is built on them, and the gap is named in the register at the foot of the page. Read them at Artificial Analysis rather than from us.

Two figures from the same publisher can be stated as they were read. The Coding Agent Index reads 67 for Astra against Claude Fable 5.1's 70, with no effort tier stated for either. Cost per index task runs $0.63 at low effort and $2.57 at max, which is 4.08 times the cost for six index points, 55 against 49: our division and our subtraction of that publisher's own per-tier figures. One index number cannot carry components moving in both directions, which is why a reader who leaves with 55 alone leaves with less than the publisher put on the page.

What Artificial Analysis does not publish is how it selects and runs those effort tiers. Its intelligence-benchmarking methodology page, read 6 September 2026, commits as a principle to "consistent prompting strategies, temperature settings, and evaluation criteria" across the models it scores. The phrase "reasoning effort" occurs on that page exactly once. Every effort setting the page fixes anywhere belongs to a judge, including a panel of three grading models named at max and high effort, and none belongs to a model under evaluation. That absence is why six numbers travel as one.

Epoch AI's detail is the most concrete capability evidence anyone has published on this model, and it is also the smallest sample on the page. On FrontierMath Erdos, Astra produced Lean-verified proofs for two of 68 open problems at $300 an attempt. Three further solutions came out of non-standardized runs costing over $220,000, which Epoch excludes from the score. Set that beside OpenAI's own summary line, which says Astra "saturates FrontierMath Tier 4 with a 98% score": the two are different sets, graded problems against open ones, and the distance between saturating the first and solving two of 68 of the second is the distance between a benchmark and a research programme. Epoch is independent as an organisation, and its FrontierMath-derived figures carry a funder relationship, because OpenAI funded FrontierMath and had visibility into most of the dataset, which Epoch disclosed itself in January 2025 after criticism. We label those figures instead of dropping them.

LMArena, read 6 September 2026, has Astra first in the WebDev/Code arena at 1797 against claude-fable-5.1-max at 1762, a 35-point gap by our subtraction. Astra does not appear on the Text, Vision, Document or Search boards. The leaderboard states no reasoning effort for Astra's entry while the entry directly beneath it carries max in its name, which is the configuration problem compressed into one row of a table.

So the direct question, whether Astra beats Claude on code, has two independent answers from the same week pointing opposite ways: ahead by 35 points on human-preference web builds, behind by three on agentic coding tasks. The tiebreak is unavailable, because no evaluator we read published the two models at equal reasoning effort for equal money. Close on both measures, and neither board on its own justifies moving a working Claude pipeline.

METR has published nothing on GPT-6 Astra as of 6 September 2026. Its blog carries no mention of Astra or GPT-6, and its most recent frontier pre-deployment write-up is GPT-5.6 Sol, dated 26 June 2026. Be precise about what that is: pre-deployment evaluations take time, and three days is evidence of nothing except that they have not published.

One third-party report belongs here because it concerns a document rather than a benchmark. ComputingForGeeks, updated 5 September 2026, reports that the model, asked directly through the API, gives its knowledge cutoff as June 2024, against the 30 April 2026 printed on OpenAI's own model page, which we read on 5 September 2026. We did not put that question to the model and have not reproduced the result. If it holds, it is a gap between a document and a self-report, and the document is the citable one.

Three further claims reached us and are not published here: a MirrorCode ranking placing Astra between two competitor models, a 100% score on something called EBR-bench, and a footnote reading "at max it hits 97.5 percent" whose referent chart we could not see. None could be verified at source on 6 September 2026, so none appears above.

What 19 people reported in the first three days

Nineteen dated reports from one platform, and the split runs along a single line: the people measuring price against measured intelligence are unimpressed, and the people running agentic and computer-use work are the opposite.

This section is Hacker News only, and that is a limitation we state rather than bury. Sixteen threads and roughly 2,900 comments were read on 6 September 2026 through the site's search API. Reddit was unreachable that day, its archive endpoint returning HTTP 429. G2, Trustpilot and Capterra do not list model APIs, so their absence is a question of scope. Where one account posted across several threads it is counted once, and every report below carries an account, a date and a link to the comment itself.

The sharpest reading arrived within hours, and it was a reading of the chart in OpenAI's own launch material. NiekvdMaas, 3 September 2026: "Title: 'major gains' / First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)". eis, the same day: "In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3." rcr-anti, also 3 September: "It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index." Fairness requires the correction that landed in the same thread: flyaway123, 3 September 2026 pointed out that the "major gains" line refers to the Artificial Analysis Coding Agent Index, which moved from 65 to 67, and not to the general index. All four were reading v4.1.1, the version live on 3 September. Three days later on v4.2 the same two models read 55 and 51, which is exactly why a version belongs beside a number.

On price the community did arithmetic in public and did not settle it. Against: simianwords, 5 September 2026 posted "GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens / GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens", concluding that for the same measured intelligence Sol costs about 40% as much. For: tosh, 5 September 2026 reported running "a few toy benches comparing Astra with Sol" and finding Astra "~30% faster and at similar cost to Sol for the same outcome", arguing that token efficiency offsets a sticker price 2.5 times Sol's. And from a buyer counting seats, MisterMunchkin, 5 September 2026: "$10/$50 is incredibly expensive compared to Chinese models which are cents ... My company is already massively cutting down on access." Other coverage of this launch benchmarks Astra's output rate against open-weight models directly; we opened none of those rate cards, so that comment is the only version of the comparison we can attribute.

Real bills, where anyone posted one. jjcm, 5 September 2026 on a single job: "That site build cost $24 - extremely non-trivial for a simple frontend." On routing, two reports worth checking against your own traffic before acting on either. d2p, 5 September 2026: "Odd that the tool call failure rate is so high (5%) for the OpenAI provider than Azure (0.2-0.5%)." And starik36, 4 September 2026 asked why anyone would use the model on Azure at all, since "It's twice as expensive, according to the link." We read no Azure rate card, so we can neither confirm nor deny that multiple, and neither failure-rate figure is one we have reproduced. What OpenAI's own pricing page does say is that its models in Amazon Bedrock are "billed through AWS and may differ from direct OpenAI pricing", so a marketplace invoice is governed by the marketplace's card and not the one above.

The positive reports cluster on one capability, and they come from people who say plainly what they did. cmrdporcupine, 6 September 2026 describes the model writing "itself a custom harness for firing up the app in different modes, taking screenshots and interacting with various screens", and concludes "It did really well." gizmodo59, 6 September 2026: "If you haven't tried computer use with Astra with codex I highly highly recommend it ... Best 200$ for an AI subscription IMHO." Those are their runs and their words. We have not run this model, and nothing on this page is a result of ours.

One pattern in the discussion itself bears on how to read every number above. Configuration disclosure is uneven, and people notice. batperson, 5 September 2026 had to ask of an untiered result, "Do you know what reasoning level this was generated at?" judge2020, 6 September 2026 wrote "I enjoy hearing that they were ran on medium effort", treating disclosure as notable. Four separate commenters stated their tier unprompted, so this is an observed failure mode and not a universal one. And aniviacat, 4 September 2026 found the case that cuts the other way, noting that "On the ScreenSpot-Pro benchmark, all effort levels achieve roughly the same score", and wondering "if that is just a limitation of the benchmark, or if the effort levels actually do not make a difference for purely visual tasks." Either way it is the right question to put to your own workload, because it is the one that tells you whether paying for max effort buys anything at all.

A last note on sources, because a reader researching this model will meet other write-ups quickly and they are not all standing in the same place. Fello AI publishes a running explainer on this launch and also sells Fello AI, a multi-model chat app, which that piece recommends as a way around an uneven access rollout; the commercial relationship appears on the page as publisher branding on a promotional box and not as a disclosure in the article body, read 6 September 2026. Yotta Labs publishes a similar piece and sells the Yotta AI Gateway, a multi-model routing product, which its piece recommends in place of defaulting to Astra; on our read of that page every outbound content link points back to its own domain and no disclosure appears in the body. Neither observation makes either piece wrong, and both contain work worth reading. It is context for weighing a recommendation. Our own position, stated so it can be held against us: this publication sells nothing adjacent to this model, and has no gateway, no routing product, no aggregator app and no affiliate arrangement with OpenAI or with any lab named on this page.

Nineteen comments were collected with an account, a date and a permalink, and all nineteen are linked above. Three days of reaction from one community is enough to show where the disagreement sits and not enough to settle it. We will re-read the rate card, the rate-limit table, both plan pages and the four evaluations whenever any of them moves, and date the change on this page when we do.

Independent evidence

Four organisations outside OpenAI had published measurable results by 6 September 2026, and every figure below carries the configuration it belongs to. ARC Prize Foundation published a full matrix on its own ARC-AGI-3 Semi-Private set, six reasoning-effort settings by two harnesses: the best Standard-harness result is 62.7% at max effort for $26,098 per task, and the best Provider Adapter result is 99.9% at high effort for $18,817, published 3 September and read 6 September. Artificial Analysis publishes six numbers for this one model on Intelligence Index v4.2, one per reasoning setting plus a non-reasoning run: 55 at max, 54 xhigh, 53 high, 52 medium, 49 low and 48 non-reasoning, with max ranked second of 202, read 6 September; the same publisher's launch article three days earlier used v4.1.1, on which Astra scored 61, and the two versions are not comparable. Epoch AI: Capability Index 169, 90% confidence interval 165 to 174, ranked first of 267, published as "Best score across settings" over seven settings rather than as one configuration, with Epoch's own note that the jump "is within our uncertainty range for the reasoning-era ECI trend"; Epoch disclosed in January 2025 that OpenAI funded FrontierMath and had visibility into most of the dataset, so its FrontierMath-derived figures carry that relationship and we label them. LMArena: first in the WebDev/Code arena at 1797 against claude-fable-5.1-max at 1762, with no reasoning effort stated for Astra's entry, read 6 September. METR has published nothing on this model as of 6 September 2026: its blog carries no mention of Astra or GPT-6, and its most recent frontier pre-deployment write-up is GPT-5.6 Sol, dated 26 June 2026.

What we read

What we did not read

  • Anthropic's rate card. We did not open Anthropic's pricing page for Claude Fable 5.1 at any point, so this review prints no per-token price for it and no Astra-against-Fable price multiple of our own. Where a reader's own comparison appears below, the arithmetic is theirs and is labelled as theirs.
  • Google's pricing and benchmark pages for Gemini 3.8 Flash. Neither was opened, so no Gemini price, no Gemini index score and no Gemini comparison of ours appears here. Separately: Gemini 3.8 Flash is named as a column on OpenAI's own announcement-page comparison table, read 3 September 2026, and our record of that table carries five values per row against six named columns, so we cannot say which value belongs to that column and we have not re-read the table to settle it.
  • What Artificial Analysis's per-component figures are measured against. Our record of that publisher's component movements for this model, read 6 September 2026, carries a direction and a size for each one and does not carry the comparison it is drawn against, whether that is the previous model, the previous index version or another effort tier, nor the reasoning setting either side of a movement ran at. We have not re-read the component pages to settle it, so those seven movements are printed below as an unresolved record and no comparison of ours is built on them.
  • Any OpenAI documentation describing an API waitlist, an application form, or the mechanics by which an account moves from one usage tier to the next. We read the published Tier 1 to Tier 5 limits on the model page and nothing that explains how an account is promoted between them, so this page states the limits and not the route.
  • Almost all of the GPT-6 Astra system card. It runs to 26,988 words and we read section 7 on hallucinations and section 4.1.1 on production benchmarks with challenging prompts. Tables 2 to 8 and Figures 1 to 17 are unread, including the under-18 evaluations, agentic safe completions, vision, static and multiturn jailbreaks, prompt injection, HealthBench, the adversarial mental-health simulations, the hallucination figures themselves, and the whole of section 8 on alignment. Specific prompt-injection and misalignment percentages circulating in other coverage of this launch sit in sections we did not open, and we do not reproduce them.
  • OpenAI's Preparedness Framework itself, and the responses API harness documentation linked from footnote 1 of the announcement page, where the two changed settings are named.
  • OpenAI's Daybreak help-centre articles and its Expanding Daybreak page, which returned 403 to us.
  • Everything behind the ChatGPT consumer pricing page. We read that page on 6 September 2026 and print its four plan prices and its Free, Plus and Pro feature lines as they appear there. The help-centre articles that set the message caps for each plan, whatever governs the advertising line on the Go plan, the Go tier's own feature list beyond that line, and any per-plan statement of which tiers reach GPT-6 Astra and on what date, are unread. That page serves Turkish from our location and a request for the English rendering redirects back, so we read the Turkish rendering only, every plan description we give from it is our translation of the Turkish, and we cannot say what the English rendering prints.
  • Whatever the string Astra belongs to on OpenAI's business pricing page. The word is present in that page's markup and does not appear in anything that rendered for us, so our finding there is about the visible plan comparison table and nothing wider. That page serves Turkish from our location too, and its English rendering is unread, so every seat description we give from it is our translation.
  • OpenAI's changelog entry for GPT-5.6 Sol's promotional price. We have the promotional rate and its stated end date from the pricing page. The earlier list price of $5 and $30 reaches us through one dated Hacker News comment that cites the changelog, and is attributed that way in the text.
  • Azure availability and Azure's rate card, and any cloud marketplace rate card. AWS Bedrock is named on OpenAI's pricing page and its billing is not published there.
  • The rate cards of the open-weight models other coverage of this launch benchmarks Astra's output price against. We read no pricing page for Qwen, DeepSeek, GLM or Kimi, so no open-weight price comparison of ours appears here.
  • Tool-call fees beyond the web-search line of $10.00 per 1,000 calls glimpsed on the pricing page.
  • The underlying API calls behind ComputingForGeeks' reported caching saving and its reported knowledge-cutoff self-report. We read the write-up, we did not reproduce either result, and we put no question to the model ourselves.
  • Reddit. The archive endpoint returned HTTP 429 on 6 September 2026, so the user evidence on this page is single-platform. G2, Trustpilot and Capterra do not list model APIs, so their absence is a matter of scope and not of access.
  • The ARC Prize and Epoch threads on x.com, which returns HTTP 402 to retrieval, and Epoch's per-benchmark pages, which render client-side with this model absent from the static HTML. Whether ARC Prize independently verified either of the two runs it publishes is unread, and so is ARC Prize's own statement of how an ARC-AGI-3 percentage is computed against its human baseline.
  • Anything published after 6 September 2026, and any document not in English except the two OpenAI plan pages named above, both of which serve Turkish from our location and were read in that rendering.

The governance and safety questions this model raises are argued separately, with the evidence, in Is GPT-6 Astra AGI? The Critical classification, and who decided who gets it. This page does not repeat that argument.

We run no hands-on tests. This review is built from the lab’s own published documents, independent evaluations by other organisations, and dated user reports, each named above with the date we read it. How we investigate →