How verdicts work
Every entry gets two separate labels. The verdict says how well the claim holds up on the evidence available. The test status says whether we've run it ourselves.
The five questions
Each claim is sent as state with five typed questions: who is making it (Choice: vendor, partner, independent), whether there's checkable evidence (Noul), whether it gives exact numbers without a method (Noul), whether comparisons were made under the same conditions (Noul), and how promotional the language is (Score, 2 to 10). The questions are published in the claim-triage recipe so you can run them yourself.
The labels
| Verified | The thing exists and the core claim checks out against a source you can open. |
| Plausible | Evidence is real but thin, or it matches a well-established pattern. Not yet reproduced. |
| Unverified | A specific claim with nothing public to check it against. |
| Misleading | The numbers may be real, but the framing or comparison makes them look better than they are. |
| Hype risk | A high-stakes use with no evidence it works as implied. |
| Toy | Real and fun, built as a joke or experiment. Not a production pattern. |
Who answers the questions
Current rubric answers are written by Claude Opus 5.5 and reviewed by an editor. We're adding Jev's answers to the same questions for every entry, and will show both. Where they disagree, that entry gets a closer look.
Tested by us
A “Tested” badge means we ran the build or recipe against Jev and Laya and recorded accuracy, latency, and cost next to the author's own figures. No entry has that badge yet. It's the part of this directory that takes the longest, and the part we think matters most.
Disputes
If you built something and think the verdict is wrong, send the evidence. Verdicts change when evidence does. Sponsorship never changes a verdict.