August 6, 2026
3 minute read. 725 words at 250 a minute, counting the opening, the section paragraphs, the call and the Monday morning list. Component notes, sources, disclosures and corrections are not counted. Scores on this page as of August 9, 2026. Any change is logged below.
VerdictHalf True
Hype Index43 / 100
SourceProceedings of the 2019 World Wide Web Conference, ACM
Call resolvesJanuary 2027
The record No call has reached its resolution date yet. The first is due January 2027. Every call is logged with a date, and the misses get published the same way the hits do.

We scored this one as holding up. Then we opened the paper we were citing, found the table, and discovered the number everybody repeats is not in it.

The claim

“Security keys stop 100% of phishing.”

Repeated in vendor material, zero-trust decks, and security awareness training since 2019. Usually without the one word in the original finding that carries all the weight.

The paper reports no targeted-attack figure for security keys at all. The famous zero comes from the blog post about the paper.

Real finding, wrong scope. Correct in part, misapplied in practice.

The verdict
Holds upAll hype
43
HALF TRUE
HYPE INDEX · 0 TO 100
What the paper prints for security keys against targeted attacks. A dash. Not zero, not a percentage. Nothing was reported for the cell this claim depends on.

What the paper actually prints

We went to the table. The famous number is not in it.

Table 3 of the paper gives protection rates for fourteen challenge types across three attacker classes: automated bots, bulk phishing, and targeted attacks. The security key row reads 100% plus or minus 25 for bots, 100% plus or minus 28 for bulk phishing, and, for targeted attacks, a dash.

A dash is not zero. It means nothing was reported for that cell. The claim in circulation is specifically about phishing that targets a person, and that is the one cell the paper leaves empty.

The widely quoted line, that zero users who used security keys exclusively fell victim to targeted phishing, comes from the company blog post summarizing the paper, not from the paper's results. We cited that blog post in our first version of this edition and repeated its figures as though they were the study's. That was the error this publication exists to catch, committed by this publication.

The margins are the story

A hundred percent with a plus or minus 28 attached is a different object from a hundred percent.

The two security key cells the paper does report are 100% plus or minus 25 and 100% plus or minus 28. Those intervals are enormous. They are what a small sample looks like when it is reported honestly, and the authors did report them honestly.

Strip the interval and you get the claim as it travels: security keys stop 100% of phishing. Keep the interval and you get something still encouraging and much more modest: in a small group, no failures were observed, and the uncertainty around that observation is wide.

By contrast the device prompt row is 100% plus or minus 1, 99% plus or minus 1, and 90% plus or minus 9, across all three attacker classes including targeted. Those are tight intervals on a large sample. The best evidenced result in the paper is not the one that became the slogan.

The word that decides whether it applies to you

Even in the blog version, one qualifier does all the work.

Exclusively. Not users who had a security key. Users for whom the key was the only way in.

A key sitting alongside an SMS fallback, a recovery code, a help-desk reset path, or a handful of exempted executive accounts has not removed the phishable step. It has added a strong option next to a weak one, and an attacker will take the weak one. The strongest authenticator sets the ceiling. The weakest one sets the floor.

The second dropped qualifier is the population. This is a study of consumer Google accounts. Applying it to an enterprise rollout is a reasonable inference and it remains an inference, because the threat model, the recovery paths, and the help desk are all different.

What holds up, and it is a lot

The verdict is about how the claim travels, not about the research.

The paper is strong work: fourteen challenge types, over 350,000 real hijacking attempts, published error margins on every cell, and a frank treatment of the friction these challenges cause. The authors are explicit that security keys offer the best immediate solution for users facing real targeted risk, and nothing here contradicts that.

Its most useful finding for anyone running a rollout is about lockout rather than efficacy. Challenges introduce friction, and the paper is candid that this locks legitimate users out. The fallback added to solve lockout is the same fallback that reopens the phishing hole.

Deploy the keys. The direction of the evidence is clear. Just do not put a hundred percent on a slide, because the paper you would be citing does not say it.

What would change our mind

We are telling you in advance what evidence would move this score.

A reported targeted-attack figure for security keys, with a margin, from any study at comparable scale, would move this well down the scale. So would an independent replication in an enterprise environment.

How this scored

Five components, twenty points each. Higher means more hype risk.

Source quality challenged3 /20
Rubric 0 to 4: Peer reviewed, or a government or regulatory dataset released with its methodology. The source is a peer-reviewed ACM WWW '19 conference paper with DOI 10.1145/3308558.3313481, placing it in the 0-4 band.
Sample and method challenged5 /20
Rubric 5 to 9: Large and well described, but the data is not available for independent inspection. Evidence describes over 350,000 real-world hijacking attempts plus 1.2M legitimate-user challenges drawn from Google login traces with method disclosed, but underlying data is not published for inspection.
Independence challenged8 /20
Rubric 5 to 9: Publisher has an indirect interest, offset by independent co-authors or external review. Co-authored by Google employees alongside NYU researchers and published via ACM; Google has commercial interest in authentication products yet the paper is not pure vendor marketing and has external academic co-authors.
Replication10 /20
Rubric 10 to 14: Not reproduced. Supporting evidence exists but comes from interested parties or a different method. Evidence contains no mention of any independent reproduction with comparable method; only the original Google/NYU results are presented.
Drift challenged17 /20
Rubric 15 to 17: Applied well outside the measured population, conditions, or timeframe. Paper states device-based challenges blocked over 94% of phishing-rooted hijacking and 100% of automated attempts, yet the circulating claim drops the automated qualifier, substitutes 'security keys', and asserts 100% against phishing.

challenged marks a component that a reader formally disputed and where the dispute was upheld, moving the score. Read the challenge and its outcome.

The full method and the rubric. If you think a component sits in the wrong band, say which band and why: challenges are published with their outcome.

Who scored this

The score above is the council's. Five models from five laboratories read the fetched text of the sources listed below, score the claim against the same published rubric independently, then review each other anonymised. Each edition is scored three times, because one run is a sample of size one, and the five component scores are the median of those runs.

Council score43 / 100, Half True
Across 3 runs47, 44, 42
Members apart6 points
Editor scored it61 / 100, Overstated

Where a person landed. Before the council took over scoring, the editor read the same sources and put this at 61, Overstated. The council came in 18 points lower, in a different band. Both readings are kept and dated, and both get graded against this edition's call when it resolves. That comparison is the point of keeping the old number rather than deleting it.

The three runs landed within 5 points of each other. Ask again and you get the same answer. The verdict held in every run.

Model Score Verdict In one line
GPT-5.6 44 Half True The evidence supports strong but imperfect protection from a broader class of device-based challenges, not a 100% phishing-prevention claim about security keys.
Claude Opus 4.8 45 Half True A real peer-reviewed finding — device-based challenges blocked 100% of *automated* hijacking and over 94% of phishing — is circulated as 'security keys stop 100% of phishing,' swapping the category the 100% actually measured.
Grok 4.5 47 Half True A real Google/NYU measurement of device-based login challenges is routinely stripped of its automated-vs-phishing distinction and re-labeled as security keys stopping all phishing.
Sonar Pro 41 Half True The original study shows 100% protection against certain measured hijacking attacks with security keys, but the popular claim that they stop 100% of phishing broadly overstates and generalizes that scoped result.
DeepSeek 45 Half True The claim overgeneralizes a finding about device-based challenges to all security keys.
Opinion of the council · written by Grok 4.5

A real Google/NYU measurement of device-based login challenges is routinely stripped of its automated-vs-phishing distinction and re-labeled as security keys stopping all phishing.

  • Source quality. The source is a peer-reviewed ACM WWW '19 conference paper with DOI 10.1145/3308558.3313481, placing it in the 0-4 band.
  • Sample and method. Evidence describes over 350,000 real-world hijacking attempts plus 1.2M legitimate-user challenges drawn from Google login traces with method disclosed, but underlying data is not published for inspection.
  • Independence. Co-authored by Google employees alongside NYU researchers and published via ACM; Google has commercial interest in authentication products yet the paper is not pure vendor marketing and has external academic co-authors.
  • Replication. Evidence contains no mention of any independent reproduction with comparable method; only the original Google/NYU results are presented.
  • Drift. Paper states device-based challenges blocked over 94% of phishing-rooted hijacking and 100% of automated attempts, yet the circulating claim drops the automated qualifier, substitutes 'security keys', and asserts 100% against phishing.

The author is not chosen. It is whichever member landed closest to the consensus.

Dissent · Sonar Pro, scored 41

One could argue for a higher Drift score because the circulating version not only drops key qualifiers but also morphs a specific dataset result into an absolute, timeless guarantee, which is much further from what the paper actually supports.

Printed verbatim. The dissenter is whichever member landed furthest from the consensus, so the strongest objection reaches you whether or not it is convenient.

Publisher's note. Nothing here is edited. The council's number is whatever the council produced, including where it disagrees with the published score and where the members disagree with each other. Agreement between models is not the same thing as accuracy, which is why the spread sits next to the median rather than under it. The scoring code is open source at github.com/marklynd/quorum, and every run writes a dated transcript before the call resolves. How the council runs, and what it is for.

The call · logged August 6, 2026 · resolves January 2027

Within the window, at least one publicly disclosed account-takeover incident at an organization with a hardware security key deployment will be attributed by that organization or by a credible independent technical analysis to a non-key authentication path that was left enabled.

How it resolves. Resolves HELD if, by January 31, 2027, at least one such incident is documented in a vendor disclosure, regulatory filing, or technical analysis by a named security research team, and the stated entry path is a non-key method such as SMS, a one-time code, a recovery flow, a help-desk reset, or an exempted account. Resolves MISSED if the only documented incidents in that window are attributed to a cryptographic or protocol defeat of the key itself. Resolves VOID if no qualifying incident is documented either way by that date.

This is a judgment call, and it goes on the record either way.

Monday morning

  1. If a deck in front of you says security keys stop 100% of phishing, ask which table that came from. The answer is that it came from a blog post about the table.
  2. Audit what is still enabled behind your keys. Every fallback, recovery flow, help-desk reset, and exemption. That list is your actual phishing exposure.
  3. Fund the lockout path before the keys arrive. An unplanned help desk becomes the weakest authenticator by default, and that is what an attacker will aim at.

Sources

Second reader. 7 of 7 printed figures traced to 4 linked sources, checked July 26, 2026. Nothing printed here is missing from the documents below. How this runs.

Every figure above traces to one of these. Where a number carries a date, use the date.

Corrections

July 26, 2026. Source changed from the Google Security Blog post to the peer-reviewed WWW 2019 paper it summarizes, which is now linked in three places including an open full text. Source quality moved from 5 to 3 and Sample and method from 11 to 8 as a result, and the total moved from 52 to 47. The headline friction figure was changed to the paper's own 52% sign-in failure rate. The blog's 38% and 34% figures remain in the body, where they are now labelled as consumer measurements reported alongside the study rather than presented as enterprise numbers. SUPERSEDED by the July 26 rewrite below, which removed that section entirely. Left here because the policy is that nothing disappears.

July 26, 2026. Independence lowered from 11 to 8. The note argued the 5 to 9 band, which covers an indirect interest offset by independent co-authors and external review, while citing 10 to 14. The band now matches the reasoning. Total moved from 47 to 44. The verdict is unchanged at Half True.

July 26, 2026. Substantially rewritten after reading the paper's Table 3 rather than the blog post summarizing it. The security key row reports 100% plus or minus 25 for bots and 100% plus or minus 28 for bulk phishing, and reports nothing at all for targeted attacks. Earlier versions of this edition repeated the blog's figures, including a recovery-phone line that does not match any row in the paper, and described a missing cell as a published zero with an unpublished denominator. Sample and method moved from 8 to 15, Independence from 8 to 15, and Drift from 13 to 16. Total moved from 44 to 61 and the verdict from Half True to Overstated. This is the most serious correction on the site and it was found by a reviewer checking our sources against the original, which is exactly the process we ask readers to run on us.

July 26, 2026. Challenge C-007. Independence moved from 8 to 15. C-005 had lowered it to 8 on the reasoning that independent co-authors and peer review put it in the 5 to 9 band. That was wrong and we upheld it wrongly. Selling the remedy a finding recommends is the 15 to 17 band, and the rubric contains no mechanic by which co-authorship pulls a direct interest below it. Peer review is why this sits at the bottom of 15 to 17 rather than the top. C-005 is marked superseded in the challenge log.

Every change is logged here rather than edited away. The policy.

Cite this

The Hype Index, August 6, 2026. “Security keys stop 100% of phishing.” scored 43/100, Half True. Primary source: Evaluating Login Challenges as a Defense Against Account Takeover, Proceedings of the 2019 World Wide Web Conference, ACM, May 2019. Scored by a council of five models, median of 3 runs, range 5; published by Mark Lynd. https://thehypeindex.com/edition/security-keys-stop-100-percent-of-phishing/

Everything in that line is checkable. That is the point of it. Think something here is wrong? Challenge the score.

Score it yourself The checker opens loaded with this edition's five component scores. Move the one you disagree with and see how much of the verdict it was carrying. It runs in your browser and nothing is sent anywhere.
Forward this
"Security keys stop 100% of phishing." — scored 43/100, Half True, by The Hype Index.
LinkedIn X Email

Every edition is public. No paywall, no registration to read.

The record

ClaimVerdictIndexCall resolves
Security keys stop 100% of phishing.
August 6, 2026
Half True 43 January 2027
There are 4.8 million unfilled cybersecurity jobs worldwide.
August 4, 2026
Overstated 76 January 2027
There is a plastic spoon's worth of microplastic in your brain, and it is causing dementia.
July 30, 2026
Half True 53 January 2027
Superhero movies are dead.
July 28, 2026
Half True 50 January 2027
95% of enterprise AI pilots fail.
July 21, 2026
Half True 53 July 2027
Get every verdict

Two editions a week. One scored claim each.

Tuesday and Thursday at 1:00 PM Central. Every call is dated and graded in public, including the ones we get wrong.

Thousands of readersFree foreverUnsubscribe anytime

One email, one scored claim, no pitch. We never sell or share your address.