Desk live·
ForensicPost
Cloud/Exposure/File 23-0918

A Single Storage Token Exposed 38TB of Microsoft AI Research Data

Researchers publishing an open-source training set generated a shareable link scoped to the whole storage account rather than to the files. Nobody attacked anything. The link was live for years and the exposure was found by an outside security firm.

Constructed geometry · not a chart of case data
TargetMicrosoft AI research
ActorExposure
D. Kennedy10 min readConfidence: high2 sources reviewed

Wiz Research published findings on 18 September 2023 describing a data exposure on a Microsoft AI GitHub repository caused by a misconfigured shared access signature token. The token, intended to share a specific set of files, was scoped to an entire storage account holding a further 38 terabytes of private data.

The exposed material was reported to include disk backups of two employee workstations, containing secrets, private keys, passwords and over 30,000 internal Teams messages from 359 employees. Wiz reported the issue to Microsoft on 22 June 2023 and the token was revoked on 24 June.

There Is No Attacker In This File

Nothing was exploited, nothing was stolen and no adversary is recorded. What happened was that a sharing mechanism offered a scope wider than the task required, and somebody accepted the default.

The corpus carries exposure as a distinct category for this reason — filed at 26-0615 for an Elasticsearch instance holding aggregated credential records. An exposure has no dwell time, no intrusion and frequently no way to know who looked.

The Grant Was Wider Than The Intent

A shared access signature is a URL that carries its own authorisation. It is convenient precisely because it needs no account, no directory and no revocation ceremony — and those same properties make it hard to inventory and hard to expire.

This is the same structural problem as the OAuth grants filed at 25-1207: an authorisation that persists until somebody actively removes it, in a system with no natural prompt to do so.

Found From Outside

The exposure was identified by an external security firm scanning for exactly this class of mistake, not by the organisation that made it. Microsoft’s investigation concluded there was no risk to customers.

The corpus records the same detection pattern at 23-1020, where affected customers spotted the Okta support breach. Where the finder is outside the organisation, the count of similar mistakes nobody was scanning for is unknowable.

How we reported this

Built on Wiz Research’s published findings, retrieved and read by this desk, and on contemporaneous reporting. The 38TB figure, the exposed content categories, the 22 June report date and the 24 June revocation are Wiz’s. Microsoft’s conclusion that there was no customer risk is Microsoft’s and is recorded as such. No individual is named and no secret, key or message content is reproduced or characterised beyond the categories Wiz published. No actor is recorded because none is alleged. Graded high. Corrections: corrections@forensicpost.com.

Sources
  1. 38TB of data accidentally exposed by Microsoft AI researchersWiz Research
  2. Microsoft AI researchers mistakenly expose 38 TB of dataTechTarget
D. Kennedy
Identity and access reporter. Former DFIR consultant. Signal on request.
// the chain of custody — tuesdays

Get the next file first.

One incident a week, taken apart properly. Logs, timelines, and what the filing left out.

PGP-signed edition · no tracking pixels · one-click unsubscribe
© 2026 ForensicPost Media · the desk · newsletter · searchGlossary