Published so you can argue with the score instead of trusting it.
A score you cannot argue with is a score you have to trust, which defeats the point. Here is what each band means on each component. Every edition cites the band it landed in.
| 0–4 | Peer reviewed, or a government or regulatory dataset released with its methodology. |
| 5–9 | Preprint, working paper, or an official statistical release that has not been reviewed. |
| 10–14 | Self-published research with a disclosed methodology, from any party. |
| 15–17 | Self-published research with no methodology, or a summary article standing in for a study. |
| 18–20 | Press release, marketing material, or no identifiable source at all. |
| 0–4 | Large, randomly selected or census-level, with the underlying data published. |
| 5–9 | Large and well described, but the data is not available for independent inspection. |
| 10–14 | Adequate size, self-selected or convenience sample, method disclosed. |
| 15–17 | Small or unrepresentative sample, or a modeled estimate presented as a measurement. |
| 18–20 | Sample size or selection not stated at all. |
| 0–4 | Published by a party with no commercial interest in the conclusion. |
| 5–9 | Publisher has an indirect interest, offset by independent co-authors or external review. |
| 10–14 | Publisher benefits from the conclusion but the finding concerns an open standard or the whole category, not its own product. |
| 15–17 | Publisher sells a product or credential positioned as the remedy for the problem the finding describes. |
| 18–20 | Publisher sells the remedy and the finding is the primary marketing asset for it. |
| 0–4 | Independently reproduced more than once with comparable method and result. |
| 5–9 | Reproduced once independently, or corroborated by a separate measure using the same method. |
| 10–14 | Not reproduced. Supporting evidence exists but comes from interested parties or a different method. |
| 15–17 | Not reproduced, and an independent measure of the same thing disagrees materially. |
| 18–20 | Not reproduced, and the original publisher has stopped standing behind it. |
| 0–4 | Used exactly as measured, on the population and conditions it was measured on. |
| 5–9 | Minor stretch. A qualifier is softened but the claim stays inside its evidence. |
| 10–14 | A load-bearing qualifier is routinely dropped, or the result is applied to an adjacent population. |
| 15–17 | Applied well outside the measured population, conditions, or timeframe. |
| 18–20 | Quoted after the source has been superseded or withdrawn, or applied to a question it never asked. |
| 0 to 19 | Underrated. The evidence is stronger than the attention it gets. |
| 20 to 39 | Holds Up. The claim survives contact with its own source. |
| 40 to 59 | Half True. Real finding, wrong scope. Correct in part, misapplied in practice. |
| 60 to 79 | Overstated. Sound evidence, carried far past what it actually measured. |
| 80 to 100 | Unsupported. No source that survives inspection. |
Every figure traces to a primary source. Not coverage of a source, the source. Where a report is quoted, the quotation comes from the report.
No number appears without somewhere to check it. If a claim cannot be sourced, it does not run.
Every edition carries a dated call that can be wrong. Calls resolve in public.
No edition ships until a second reader has opened every linked primary source and matched every figure in the edition against it, line by line. Not the abstract, not the press release, not the blog post about the paper: the table the number sits in.
This exists because we failed it. The security keys edition originally quoted figures from a company blog post as though they were the peer-reviewed paper's results, and one of them disagreed with the paper by 73 points. A build-time validator caught none of it, because arithmetic that is internally consistent can still be copied from the wrong document.
Who runs it. An independent reviewer, briefed to find problems rather than to confirm quality, and paid whether or not they find any. What it produces. Every figure the check moves becomes a published correction and, where it started as a challenge, a numbered entry in the challenge log.
What it produced, this run. There are 4.8 million unfilled cybersecurity jobs worldwide (August 4, 2026): 22 of 22 figures traced to 6 sources. 95% of enterprise AI pilots fail (July 21, 2026): 21 of 21 figures traced to 3 sources. There is a plastic spoon's worth of microplastic in your brain, and it is causing dementia (July 30, 2026): 9 of 9 figures traced to 9 sources. Security keys stop 100% of phishing (August 6, 2026): 7 of 7 figures traced to 4 sources. Superhero movies are dead (July 28, 2026): 17 of 17 figures traced to 9 sources. Checked July 26, 2026. The script is in the repository and it exits non-zero if a figure cannot be traced, which blocks the deploy.
What it is not. It is not a person, and it is not independent of this publication. It reads the linked documents rather than the summaries of them, which is the specific failure that produced our worst correction, and it is run on a separate pass from the scoring with a brief to find problems. Those are real properties and they are narrower than editorial independence. A second human who is not the editor would be better, and we do not have one.
One person scoring every edition was a real weakness, and disclosure did not fix it. Since 27 July 2026 the score is produced by a council: five flagship models from five different laboratories, reading the fetched text of the primary sources and scoring the same claim against this same published rubric. The editor publishes what the council produces and does not move it.
What changed, and what it cost. Every edition scored before that date was rescored. The council came in lower on all five, by between three and eighteen points, and moved the verdict band on three of them. Where an edition had already gone out to subscribers the change is written into its corrections log, because a number a reader has already seen does not get altered quietly. The editor's original score is kept on every edition page. When these calls resolve in January, both readings get graded against the same outcome, which is the entire reason for keeping the old number instead of deleting it.
How a run works. Each model is handed the claim, the rubric, and the fetched text of the primary sources, and is told to score from that text rather than from what it remembers. Members never see who else is on the panel: identities are replaced with letters and the order is shuffled on every call, so a model cannot defer to a brand. The opinion of the council is written by whichever member landed closest to the consensus, and the dissent by whichever landed furthest. Neither is chosen by a person.
Why three runs. One council run is a sample of size one. The same claim, the same evidence and the same rubric can produce different totals on different runs, so every edition is scored three times and the median is published with the observed range beside it. A claim whose score holds within a point and a claim whose score swings twenty points are different objects, and you get to see which one you are looking at.
What a failure looks like. A model that times out, returns something unparseable, or scores outside the rubric is recorded as a failure with its reason. It is never replaced with a default or a mid-range guess. If fewer than three members return a usable score, no consensus is reported at all.
The honest limitation. Agreement between models is not accuracy. These systems train on overlapping material, so they can be wrong together and the agreement will make the error more persuasive rather than less. That is exactly why the spread sits next to the median rather than under it, and why the members' individual scores are printed even when they embarrass the consensus.
The scoring library is open source and MIT licensed at github.com/marklynd/quorum, including the rubric loader, the anonymised review, and the tests. Every run writes a dated transcript with a SHA-256 hash of each round, recorded before the call resolves, so when a call lands held or missed it can be shown how each model scored it in advance.
A council is still not a guarantee, and a rubric applied consistently can be applied consistently wrong. So there is a process instead of a promise.
Anyone can dispute a component. Name the edition, name the component, and say which published band you think it belongs in and why. Send it to [email protected]. You do not need credentials and you do not need to be right.
Every challenge gets a published outcome. Upheld, partly upheld, or declined, with the reasoning. An upheld challenge is a correction and moves the score. A declined one is published too, with the argument against it, so you can judge whether the decline was fair.
Challenged scores are marked. A component that has survived a specific challenge carries more weight than one nobody has tested, and the record will say which is which.
The first challenges came from a hostile review commissioned before launch, with a brief to find problems rather than confirm quality. Every one it raised was upheld and every one moved a score. We are publishing them because a process nobody has run is a promise, and because the failures it found are the strongest evidence the method works on its own author.
Challenge. The edition scored Source quality on the strength of peer-reviewed papers while linking only a company blog post summarizing them, and the two paper links returned 404.
Outcome. Correct and serious. The links were dead. The edition now cites the open full text hosted by a co-author, the ACM Digital Library entry as the citation of record, the publisher's research listing, and the blog summary, in that order. Source quality moved from 5 to 3 and Sample and method from 11 to 8. Edition total 52 to 47.
Challenge. Source quality scored 15, citing the 15 to 17 band for research with no disclosed methodology, while the note stated the 2024 study disclosed its methodology.
Outcome. Correct. Moved from 15 to 13, the 10 to 14 band. The withdrawal of the metric is scored at the top band on Replication and Drift, where it belongs, rather than counted twice. Edition total 81 to 79, and the verdict moved from Unsupported to Overstated.
Challenge. Sample and method scored 12 in the 10 to 14 band, while the note itself described a gap modeled from OECD and BLS baselines. The published rubric puts a modeled estimate presented as a measurement at 15 to 17.
Outcome. Correct. Moved from 12 to 15. Edition total 79 to 82, which returned the verdict to Unsupported.
Challenge. Sample and method scored 14 and was justified as adequate size for 52 interviews, while the same edition's headline statistic presents 52 as the problem.
Outcome. Correct, and an embarrassing one. Fifty-two interviews is a small sample for a claim about enterprise AI generally. Moved from 14 to 16, the 15 to 17 band. Edition total 67 to 69. The verdict is unchanged at Overstated.
Challenge. Independence was scored 11, citing the 10 to 14 band, while the note's own justification named two independent university co-authors and peer review, which is the defining language of the 5 to 9 band.
Outcome. Originally upheld, then reversed. See C-007. Selling the remedy a finding recommends is the 15 to 17 band, and the rubric contains no mechanic by which independent co-authorship pulls a direct commercial interest below it. Lowering Independence to 8 was wrong, and upholding the challenge that argued for it was wrong. Both are left published.
Challenge. The edition presented figures from the company blog post as the paper's published protection rates, including a recovery-phone line that matches no row in the paper. It also described the security key result as a published zero with an unpublished denominator, when the paper reports no targeted-attack figure for security keys at all and publishes plus or minus 25 and plus or minus 28 margins on the two cells it does report.
Outcome. Correct, and the most serious finding raised against this publication. The edition was rewritten against Table 3 of the paper. Sample and method moved from 8 to 15, Independence from 8 to 15, and Drift from 13 to 16. Edition total 44 to 61, and the verdict moved from Half True to Overstated. Source quality is unchanged at 3: the research is strong, and the failure was ours in citing a summary of it.
Challenge. C-005 was upheld on the reasoning that two independent university co-authors and peer review put Independence in the 5 to 9 band. But the same note also stated that the publisher sells security keys and benefits from the conclusion, which is the defining language of 15 to 17. The log moved the score without ever reconciling the two, and a later entry moved it back with no stated reason.
Outcome. Upheld against ourselves. Independence is 15, the 15 to 17 band, because the publisher sells the remedy the finding recommends. Peer review and independent co-authors are why it sits at the bottom of that band rather than the top, not why it leaves it. C-005 is now marked Declined, with its original argument left published. Edition total 61 to 61: the component was already at 15 when the edition was rewritten, and this entry supplies the reasoning that was missing.
A publication that scores other people's evidence has to be harder on its own. Three rules.
Errors are corrected in place and logged, never quietly edited. If a figure, a source, or a component score is wrong, the edition is updated and a dated correction note is added to it, stating what changed and why. Superseded notes stay published and are marked superseded rather than deleted.
What we do not yet do. We do not archive a full copy of every prior version of an edition, so where a rewrite was substantial you have the correction note rather than a diff. That is a real gap and it is stated here rather than papered over. The security keys edition was rewritten substantially on July 26, 2026 and is the one case where it matters today.
A changed score is a correction, not a revision. If new evidence moves a component, the score changes and both the old and new numbers are shown with the reason. That is the system working, and hiding it would defeat the purpose of publishing components at all.
Calls resolve on their stated date and stay on the record. Held or missed, the outcome is published the same way. A record that only shows the hits is marketing.
Found something wrong, or think a component is mis-scored? Write to [email protected]. Specific disagreements are more useful than agreement, and the good ones get published.
A score from zero to one hundred applied to a public claim. Higher means more distance between what the evidence supports and how the claim is being used. It is the sum of five components worth twenty points each.
Most fact checking asks whether a study is right. It rarely asks how far the retelling has carried the finding past what it measured. That gap is where people get hurt, so it is scored on its own.
No. A claim can rest on solid research and still score high because it is being used far outside its scope. That is the difference between Overstated and Unsupported.
It is marked held or missed on the record and it stays there. The archive is the point.
Errors are corrected in place and logged with a dated note, never quietly edited. If new evidence changes a component score, both the old and new numbers are published with the reason. Calls resolve on their stated date and stay on the record whether they held or missed.
Tuesday and Thursday at 1:00 PM Central. Every call is dated and graded in public, including the ones we get wrong.
One email, one scored claim, no pitch. We never sell or share your address.