AI Tools Police
Reader-supported — we may earn a commission from links, at no cost to you. Rankings are never sold. How we investigate →

How we investigate

This page is the standard we hold ourselves to. It sets out what we verify before a review goes live, how we label the evidence behind every claim, how scores are decided, and what we refuse to do. If a page on this site breaks one of these rules, tell us and we will fix it.

What we verify

Every tool is judged on the same six dimensions, whatever the category. The subscores change to fit the product. These do not.

Verified pricing

Plans, credit limits, and renewal prices are checked against the vendor's own live pricing page, and the review carries the date that check happened. Prices in this market move without notice, so a dated number tells you exactly how fresh it is. When a review shows no verification date, the page does not record a dated check of the vendor's own pricing page, and the figures should be read as coming from documentation, an archived capture, or a source the review names. We would rather show you the gap than paper over it. Where a plan hides its real cost inside credits or seats, we work out what the tool costs to actually use.

Documented features

Capabilities come from official documentation, changelogs, and the vendor's own support pages, not from marketing adjectives. If a feature is gated behind a higher tier, still in beta, or capped, we say which. A capability that exists only in a launch announcement is reported as an announcement.

Real user reports

We aggregate first-hand experience from G2, Trustpilot, Reddit, and Capterra, then weigh recurring complaints and praise against the documented feature set. One angry review is noise. The same complaint repeating across independent platforms is a pattern, and patterns go in the review.

Transparent trade-offs

We name the exact point where a free tier stops being useful, where a competitor does the job better, and what the real risks are. A page that lists only upside is an advertisement, not a review.

Honest-ceiling scoring

Scores come from evidence, and evidence caps them. No tool here is perfect, so a top pick usually lands in the mid-to-high range rather than at a flawless number. Where the record on a dimension is thin, the score for it stays conservative. The headline score is not the average of the dimension scores beneath it, and you should not expect the arithmetic to work: the dimensions are separate judgements about separate things, and the headline says what the tool is worth to the reader it suits. A detector can score 4.0 on raw AI text and 2.0 on writing by non-native speakers, and neither number is the answer on its own. Read the dimensions, because that is where the disagreement with the headline lives, and it is usually the more useful number for your case.

New insight only

A review goes live only if it adds something the existing results do not already cover. We do not publish a reworded version of what is already out there, and we do not publish to fill a calendar.

How we label evidence

We do not run hands-on tests, and we never write as though we did. Our method is verification: official documents, live pricing pages, and aggregated independent user reports. The strength of that method is that anything we publish can be traced back to a source you can open yourself; the limit is that we cannot tell you how a tool felt to use for a month. We state both rather than blurring them. Every substantive claim in a review carries one of five evidence classes.

LabelWhat it meansExample phrasing
VerifiedConfirmed against an official document, live pricing page, or policy“The Business plan includes one editor seat (verified on the vendor's pricing page, July 2026).”
User-reportedReported across multiple independent user sources such as G2, Trustpilot, Reddit, and Capterra“Multiple users report billing continuing after cancellation.”
Vendor-reportedThe company's own claim, with no independent confirmation“The company claims 99% accuracy; we found no independent confirmation.”
EstimatedOur own calculation from stated assumptions“At the published credit rate, 100 videos would cost roughly $X.”
UnknownCould not be verified, or the company does not disclose it“The refund policy for failed generations is not documented.”

There is no sixth label covering first-hand product trials, because we do not conduct them. A claim verified from documents and user reports is presented as exactly that, and never dressed up as personal experience.

Where our facts come from

Not every source carries the same weight, so we grade them before we use them.

TierSourcesHow we use them
A. PrimaryOfficial documentation, live pricing pages, legal texts, filings, academic papersThe basis for any critical claim
B. Reputable secondaryEstablished independent press, expert interviews, verified benchmarksContext and corroboration
C. User signalG2, Trustpilot, Reddit, Capterra, community forumsExperience patterns, weighed but never treated as settled fact on their own
D. PromotionalAffiliate blogs, aggregators, marketing listiclesDiscovery only, never a final source

A primary source beats a secondary one whenever both exist. A statement from the company is attributed to the company, not presented as independent fact. Where a claim is contested, we label it contested and show the disagreement rather than picking the version that reads better. Our full citation standard is on the sources page.

How a review is made

1

Research

We map the tool, its category, its alternatives, and the questions people are actually asking about it, including the ones competing reviews leave unanswered.

2

Draft

We write the verdict first: who the tool is for, where it wins, where it breaks. The summary sits at the top of the page and has to stand on its own for a reader who never scrolls.

3

Fact-check

Every number is checked before it goes live. Pricing against the vendor's live page, features against the documentation, claims against independent sources. If something cannot be confirmed, it is either labeled as unverified or it does not make the page.

4

Publish and maintain

The review goes live dated and bylined, and we revisit it as pricing, features, and company claims change. Reviews are snapshots of a moving market, and we treat them that way.

How scoring works

Every review of a tool still in operation carries one overall score out of five, plus four to six subscores chosen to fit the category. A headshot generator is scored on likeness and keeper rate. A voice tool is scored on dimensions like naturalism, cloning, or language coverage. Forcing every category through the same five labels would flatter tools that are weak at the things that matter most in their own market.

Two dimensions are weighed in every category, whatever the product does. Economic reality covers the true cost of ownership: credits, seats, renewal pricing, and commercial rights. Trust covers billing behavior, privacy and data handling, support responsiveness, and whether the company's own claims survive scrutiny.

Scores are capped by evidence. Where the record is thin, or a company's central claim has no independent confirmation, the score reflects that instead of giving the benefit of the doubt. No tool here carries a perfect overall score. A tool can be the best pick in its category and still carry a number that tells you to read the trade-offs before you pay.

How reviews stay current

Every review is dated. The original publication date is recorded and never changes, because moving it would tell you the page is newer than the work behind it. The "last updated" date changes only when something material changed: a price, a feature, a limit, a policy, or the verdict itself. Cosmetic edits do not reset it.

When we revisit a review, pricing is rechecked against the vendor's live page rather than against our own previous version. Where a review carries a pricing table, that table shows the date the numbers were last checked. When a price turns out to be wrong, we correct it everywhere it appears on the site, including on pages that compare a competitor against it.

What we don't do

A company can report a factual error after publication, the same as any reader. That is the whole of its access.

Independence

Rankings are never sold. Affiliate commission does not move a score, change a position, or decide whether a tool gets covered at all, and tools without an affiliate program are reviewed on the same terms as tools that pay. Paid options exist, they are published openly, and none of them buys a score or a place in a ranking. The full policy is on our affiliate standards page.

Corrections

We fix errors in public and we date the fix. Corrections are graded by how much they affect the reader, from a silent typo fix to a prominent note and a revised verdict, and each grade carries a response time we hold ourselves to. Those levels and target times are set out on our corrections page.

info@aitoolspolice.com