We scored this one as holding up. Then we opened the paper we were citing, found the table, and discovered the number everybody repeats is not in it.
The paper reports no targeted-attack figure for security keys at all. The famous zero comes from the blog post about the paper.
We went to the table. The famous number is not in it.
Table 3 of the paper gives protection rates for fourteen challenge types across three attacker classes: automated bots, bulk phishing, and targeted attacks. The security key row reads 100% plus or minus 25 for bots, 100% plus or minus 28 for bulk phishing, and, for targeted attacks, a dash.
A dash is not zero. It means nothing was reported for that cell. The claim in circulation is specifically about phishing that targets a person, and that is the one cell the paper leaves empty.
The widely quoted line, that zero users who used security keys exclusively fell victim to targeted phishing, comes from the company blog post summarizing the paper, not from the paper's results. We cited that blog post in our first version of this edition and repeated its figures as though they were the study's. That was the error this publication exists to catch, committed by this publication.
A hundred percent with a plus or minus 28 attached is a different object from a hundred percent.
The two security key cells the paper does report are 100% plus or minus 25 and 100% plus or minus 28. Those intervals are enormous. They are what a small sample looks like when it is reported honestly, and the authors did report them honestly.
Strip the interval and you get the claim as it travels: security keys stop 100% of phishing. Keep the interval and you get something still encouraging and much more modest: in a small group, no failures were observed, and the uncertainty around that observation is wide.
By contrast the device prompt row is 100% plus or minus 1, 99% plus or minus 1, and 90% plus or minus 9, across all three attacker classes including targeted. Those are tight intervals on a large sample. The best evidenced result in the paper is not the one that became the slogan.
Even in the blog version, one qualifier does all the work.
Exclusively. Not users who had a security key. Users for whom the key was the only way in.
A key sitting alongside an SMS fallback, a recovery code, a help-desk reset path, or a handful of exempted executive accounts has not removed the phishable step. It has added a strong option next to a weak one, and an attacker will take the weak one. The strongest authenticator sets the ceiling. The weakest one sets the floor.
The second dropped qualifier is the population. This is a study of consumer Google accounts. Applying it to an enterprise rollout is a reasonable inference and it remains an inference, because the threat model, the recovery paths, and the help desk are all different.
The verdict is about how the claim travels, not about the research.
The paper is strong work: fourteen challenge types, over 350,000 real hijacking attempts, published error margins on every cell, and a frank treatment of the friction these challenges cause. The authors are explicit that security keys offer the best immediate solution for users facing real targeted risk, and nothing here contradicts that.
Its most useful finding for anyone running a rollout is about lockout rather than efficacy. Challenges introduce friction, and the paper is candid that this locks legitimate users out. The fallback added to solve lockout is the same fallback that reopens the phishing hole.
Deploy the keys. The direction of the evidence is clear. Just do not put a hundred percent on a slide, because the paper you would be citing does not say it.
We are telling you in advance what evidence would move this score.
A reported targeted-attack figure for security keys, with a margin, from any study at comparable scale, would move this well down the scale. So would an independent replication in an enterprise environment.
Five components, twenty points each. Higher means more hype risk.
The score above is the council's. Five models from five laboratories read the fetched text of the sources listed below, score the claim against the same published rubric independently, then review each other anonymised. Each edition is scored three times, because one run is a sample of size one, and the five component scores are the median of those runs.
| Model | Score | Verdict | In one line |
|---|---|---|---|
| GPT-5.6 | 44 | Half True | The evidence supports strong but imperfect protection from a broader class of device-based challenges, not a 100% phishing-prevention claim about security keys. |
| Claude Opus 4.8 | 45 | Half True | A real peer-reviewed finding — device-based challenges blocked 100% of *automated* hijacking and over 94% of phishing — is circulated as 'security keys stop 100% of phishing,' swapping the category the 100% actually measured. |
| Grok 4.5 | 47 | Half True | A real Google/NYU measurement of device-based login challenges is routinely stripped of its automated-vs-phishing distinction and re-labeled as security keys stopping all phishing. |
| Sonar Pro | 41 | Half True | The original study shows 100% protection against certain measured hijacking attacks with security keys, but the popular claim that they stop 100% of phishing broadly overstates and generalizes that scoped result. |
| DeepSeek | 45 | Half True | The claim overgeneralizes a finding about device-based challenges to all security keys. |
A real Google/NYU measurement of device-based login challenges is routinely stripped of its automated-vs-phishing distinction and re-labeled as security keys stopping all phishing.
One could argue for a higher Drift score because the circulating version not only drops key qualifiers but also morphs a specific dataset result into an absolute, timeless guarantee, which is much further from what the paper actually supports.
Within the window, at least one publicly disclosed account-takeover incident at an organization with a hardware security key deployment will be attributed by that organization or by a credible independent technical analysis to a non-key authentication path that was left enabled.
July 26, 2026. Source changed from the Google Security Blog post to the peer-reviewed WWW 2019 paper it summarizes, which is now linked in three places including an open full text. Source quality moved from 5 to 3 and Sample and method from 11 to 8 as a result, and the total moved from 52 to 47. The headline friction figure was changed to the paper's own 52% sign-in failure rate. The blog's 38% and 34% figures remain in the body, where they are now labelled as consumer measurements reported alongside the study rather than presented as enterprise numbers. SUPERSEDED by the July 26 rewrite below, which removed that section entirely. Left here because the policy is that nothing disappears.
July 26, 2026. Independence lowered from 11 to 8. The note argued the 5 to 9 band, which covers an indirect interest offset by independent co-authors and external review, while citing 10 to 14. The band now matches the reasoning. Total moved from 47 to 44. The verdict is unchanged at Half True.
July 26, 2026. Substantially rewritten after reading the paper's Table 3 rather than the blog post summarizing it. The security key row reports 100% plus or minus 25 for bots and 100% plus or minus 28 for bulk phishing, and reports nothing at all for targeted attacks. Earlier versions of this edition repeated the blog's figures, including a recovery-phone line that does not match any row in the paper, and described a missing cell as a published zero with an unpublished denominator. Sample and method moved from 8 to 15, Independence from 8 to 15, and Drift from 13 to 16. Total moved from 44 to 61 and the verdict from Half True to Overstated. This is the most serious correction on the site and it was found by a reviewer checking our sources against the original, which is exactly the process we ask readers to run on us.
July 26, 2026. Challenge C-007. Independence moved from 8 to 15. C-005 had lowered it to 8 on the reasoning that independent co-authors and peer review put it in the 5 to 9 band. That was wrong and we upheld it wrongly. Selling the remedy a finding recommends is the 15 to 17 band, and the rubric contains no mechanic by which co-authorship pulls a direct interest below it. Peer review is why this sits at the bottom of 15 to 17 rather than the top. C-005 is marked superseded in the challenge log.
The Hype Index, August 6, 2026. “Security keys stop 100% of phishing.” scored 43/100, Half True. Primary source: Evaluating Login Challenges as a Defense Against Account Takeover, Proceedings of the 2019 World Wide Web Conference, ACM, May 2019. Scored by a council of five models, median of 3 runs, range 5; published by Mark Lynd. https://thehypeindex.com/edition/security-keys-stop-100-percent-of-phishing/
| Claim | Verdict | Index | Call resolves |
|---|---|---|---|
| Security keys stop 100% of phishing. August 6, 2026 |
Half True | 43● | January 2027 |
| There are 4.8 million unfilled cybersecurity jobs worldwide. August 4, 2026 |
Overstated | 76● | January 2027 |
| There is a plastic spoon's worth of microplastic in your brain, and it is causing dementia. July 30, 2026 |
Half True | 53 | January 2027 |
| Superhero movies are dead. July 28, 2026 |
Half True | 50 | January 2027 |
| 95% of enterprise AI pilots fail. July 21, 2026 |
Half True | 53● | July 2027 |
Tuesday and Thursday at 1:00 PM Central. Every call is dated and graded in public, including the ones we get wrong.
One email, one scored claim, no pitch. We never sell or share your address.