Within days of the 16-billion credential figure filed at 25-0620 circulating, other analysts pushed back. Their argument: the trove consisted largely of recycled and outdated data, repackaged from prior leaks, with some portion potentially fabricated — and therefore far less significant than the headline suggested.
We publish both, because the disagreement is the most useful thing about the episode.
Nobody Was Lying
The researchers who found the datasets counted records and reported a count. The analysts who disputed it asked how many described a currently-valid credential for a real person.
Those are different questions with different answers, and neither team was wrong about its own. It is the same measurement problem this desk filed at ADT in 26-0425 and 7-Eleven in 26-0423: rows counted versus people affected.
The Incentive Structure Produces The Big Number
A record count is available immediately, is objectively checkable, and produces coverage. A unique-and-valid count requires deduplication against prior corpora and validation work that takes weeks, by which point the story has moved on.
So the first number published is almost always the largest one, and the correction — where it comes — reaches a fraction of the original audience.
What This Desk Does With It
We publish the claim, name the claimant, publish the dispute, name the disputer, and grade the file on what is established rather than on what is asserted.
It is the same standard applied to attacker claims throughout this database. A researcher’s figure is not an attacker’s figure, and it still benefits from being labelled as a measurement made by a particular party using a particular method.
Compiled from competing published analyses, listed below. We take no position on which estimate is closer to correct; neither has been independently reconciled. Corrections: corrections@forensicpost.com.