How to read a launch-week claim
Updated 23 Sep 2026
New model categories attract big numbers fast. Most of them aren't lies. They're true numbers in a frame that makes them look more general than they are. These are the five checks we run on every entry.
Who is saying it?
A vendor, a paid partner, and an unaffiliated builder can all be right, but they have different reasons to round up.
Can you open the evidence?
A repo, a live demo, or a dataset someone else can run. A screenshot of a result isn't evidence of the result.
Do the numbers show their method?
“23 turns and 39.6 seconds” sounds rigorous. Without the task, the hardware, and the number of runs, precision is decoration.
Is the comparison fair?
The most common problem in launch week: comparing local GPU inference with a hosted API's network round trip, or quoting a score from a model fine-tuned on the benchmark's own training data.
How promotional is the language?
“Changes everything” and “can't hallucinate” are claims too, and they rarely survive the first four checks.
See how verdicts work for how these five answers become a label.