A dataset covering around 2.6 million Duolingo accounts was published on a hacking forum in August 2023, having previously been offered for sale in January. Reporting describes it as scraped through an application programming interface that returns profile information for a supplied username.
Duolingo was reported as saying the data was scraped from public profile information. The published dataset also contained email addresses, which are not part of a public profile.
Why This Is Graded Medium
What is established is that the dataset exists, that it was published, and that an interface of the described kind was openly reachable. What is not established is the boundary between what the interface returned and what the compiler already held.
The company’s characterisation and the dataset’s contents do not fully reconcile, and this desk has not seen the data. Grading it high would assert a resolution nobody published; grading it low would ignore consistent reporting and a company statement. Medium is the honest position.
Scraping Is Not A Lesser Category
There is a persistent argument that scraped data is not a breach because each field was individually obtainable. The corpus rejects the framing on the same grounds it applied at 26-0615: aggregation is the harm.
A profile visible to someone who already knows a username is a different object from 2.6 million profiles in a downloadable file, and the second one was never available to anybody until an interface made it cheap to assemble.
Confirmation Is The Dangerous Feature
The reported capability that matters most is that an email address could be submitted and confirmed as belonging to an account. That turns a list of addresses from any other breach into a targeting list for this service.
The corpus argues at 23-0119 that an API is a door built for machines. This is the second-order version: an interface that answers yes or no about a person is a service for whoever asks it a few million times.
Built on contemporaneous reporting of the dataset’s publication and of the scraping method. The 2.6 million figure, the January sale and August republication, and the description of the API are from that reporting. Duolingo’s characterisation of the data as public profile information is the company’s, reported at the time; this desk notes that it does not obviously account for the email addresses in the dataset and has not resolved the discrepancy. This desk has not obtained or examined the dataset. Graded medium: the existence and mechanism are well reported, the scope is not established. No sample data is reproduced. Corrections: corrections@forensicpost.com.
- Scraped data of 2.6 million Duolingo users released on hacking forumBleepingComputer
- Data of 2.6 Million Duolingo Users Leaked on Hacking ForumInfosecurity Magazine