Planned
What a new frontier model actually does, and what it actually costs
When a major lab ships a model, the pages that appear within hours are mostly the lab's own numbers rewritten. This section is planned for the other job: reading the system card, the rate card and the independent evaluations, and saying plainly what the model is for, what it costs to run, and which of its headline figures the lab itself says should not be read the way everyone reads them.
What this section will be
- — Every review will lead with the model's own name, because that is what a reader searches for, and will state a verdict in one sentence you can disagree with.
- — Every figure will say whose it is. A lab's own benchmark is a vendor claim, not an independent result, and the two will never be printed as though they were the same thing.
- — Every benchmark score will name the configuration it belongs to. A reasoning-effort tier or a special access level changes the number, and a score without its configuration is an incomplete claim.
- — Every price will carry its billing basis and its multipliers. A per-token headline rate is not the bill when caching, batch, fast mode and long-context surcharges move it in both directions.
- — Every review will list the documents we read with the date we read each one, and will list what we did not read. These documents run to tens of thousands of words and we will not pretend otherwise.
- — Where a model raises a governance or safety question, we will link to the analysis that argues it rather than re-summarising it here.
What you can read today
Live today are 77 dated investigations of AI products, an original pricing-transparency study with an open dataset, and an analysis desk that has already read one frontier launch against the EU AI Act. This section is the same discipline pointed at the models themselves.