In a classic fairy tale trope, the villagers give a dragon a beautiful maiden every year in exchange for a quiet life. Imagine a story where the dragon ate all the maidens, devoured the people who brought them, then finished off the village and started eating up the graveyard.
On August 14, 2026, it emerged that Google was buying the entire internal archive of a bankrupt American airline.
A fairy tale dragon could take a hundred years to start feasting on the dead. AI models are hungrier: they made it to the graveyard in just four.
For free
Data scraping started in an age of abundance. Foundation models grew up on the open web, scraped into giant corpora like Common Crawl and C4. Everything that lay unlocked went in: encyclopedias, news, forums, washing-machine manuals, court rulings, and strangers arguing about the right way to cook rice.
All the information in those archives has one thing in common — it was written to be seen. A person leaving a comment or posting an article knows that outsiders are watching. They filter their style and edit their thoughts. The open web is a display window, and that's exactly what the models learned: how people perform for an audience, even a casual one. It almost entirely missed how people actually talk to each other.
The age of abundance ended abruptly, and not because the internet ran dry. In 2024, researchers at the Data Provenance Initiative checked 14,000 domains used to train AI. It turned out that in just one year, 28% of the most important sources had banned data collection, and 45% of the C4 dataset had come under new restrictions.
Web users realized they weren't just an audience — they were an asset — and organizations started shutting their doors accordingly.

On lease
A second era had begun: the era of contracts.
Reddit’s IPO prospectus, in March 2024, disclosed data-licensing contracts worth $203 million. This sum included a $60 million annual agreement with Google, which the company had signed a month before going public. Three months later, Reddit and OpenAI agreed on a deal worth roughly $70 million per year.
The community platform had opened the floodgates. By the summer of 2026, more than 90 such deals had been signed, including a five-year, $250 million agreement between News Corp and OpenAI. Wiley offered its data (giving authors no option to opt out) for $44 million, while Taylor & Francis’s deal with Microsoft had an £8 million value for its first year (and authors claim they found out about it in the news).
The most transparent and the rarest agreement is that between HarperCollins and Microsoft: $5,000 per book, split evenly between publisher and author, participation strictly opt-in, no more than 200 consecutive words or 5% of the text, fiction excluded. If this sort of contract becomes common, it will show that the industry is growing up and learning from multiple lawsuits.
Reddit brought such a lawsuit against Anthropic, the company that talks louder than anyone about safety and respect for rights holders. Reddit claimed that its bots hit the platform more than 100,000 times without paying for access. After Anthropic tried and failed to move the case to federal court, it became a bookkeeping exercise over a commercial contract: which company owed the other and how much. Ever since, when you ask it to look up something interesting on Reddit, Claude Fable replies something like this:
The dragon has been locked out of the front entrance, so now it politely asks a human to hold the door. So — connect the extension [Chrome] and I'll gather the threads for you.
Reddit blocks the browser extension just fine, by the way. The model simply has a hard time grasping that.
All these deals share three features in common.
First, they’re leases, not sales. The data stays with the owner; the buyer gets the right to use it under certain conditions for a certain time.
Second, the seller holds leverage. It can always raise the price or refuse to renew the license. After the AI-generated summaries in Google’s search results started eating into the Reddit click-through activity the companies’ deal was supposed to deliver, Reddit started debating whether or not to renew the agreement. As of this writing, it hasn’t publicly reached a decision.
Third, and most importantly, the models are feasting on the same corpus as before. Reddit sells only public posts. Publishers sell only published articles. Putting them behind a paywall hasn’t taken LLMs beyond the display window, just charged them to look at all.

For good
The Internet also contains a vast layer of text that was never meant for outside readers at all, like internal work correspondence. Decades of archives show how people actually negotiate: who gives ground and on which word, what a refusal dressed up as agreement looks like, and how many emails it takes to settle something that could be decided in a minute. None of it was edited, none written with an eye toward posterity.
An operating company will never hand its archive over — it's the corporation's lifeblood, its reputation, its biggest legal exposure.
Until now, the public only had access to one such archive: that of Enron.
The energy corporation collapsed amidst scandal in 2001, and regulators seized its servers. A year and a half later, they released the company’s internal correspondence, which contained roughly 500,000 emails from about 150 employees. Researchers at MIT bought the archive for $10,000, and in 2004, Carnegie Mellon University turned it into the legendary Enron Email Dataset. Over more than 20 years, 20,000 academic papers have cited that data. It taught algorithms to catch spam, to build relationship graphs, to recognize emotion, and to reconstruct chains of dialogue. In 2025, the archive became the basis of EnronQA, the leading benchmark for testing how well neural networks read other people's corporate secrets.
Modern AI learned almost everything it knows about workplace communication from the emails of a single bankrupt Texas company in the early 2000s. The material wasn't uniquely good, simply all that was available.
CMU has cleaned up the dataset repeatedly, removing some emails at the request of the people involved, and released a corrected version in 2015. Today the archive exists in multiple forms: a research version, a cleaned-up version, and an annotated version. Even if someone involved with the original release suddenly decided publication had been a mistake, there's no undoing it now. The archive stopped being an object long ago and became an environment.

The inheritance of the dead
For more than 20 years, Enron’s data was unique. An investigation had exposed correspondence and researchers had taken the emails public. Situations like that simply don’t come along often, and nobody had tried to specifically acquire company data for AI training — not until spring 2025, when DNA-testing company 23andMe went bankrupt and the data of its 15 million customers went up for auction.
Because the data in question was genetic, the government stepped in: attorneys general from 27 states opposed selling the 23andMe database without consent from its users. The Department of Justice appointed an independent consumer privacy ombudsman, an action it could take because the state has a mandate to protect customers.
When Spirit Airlines declared bankruptcy in 2026, it sold its assets piece by piece, becoming the first company of its size to do so in 25 years. One of those assets was a type that had never appeared on an auction block before: employees' correspondence, all their work files, and the corresponding technology.
In the past, dead companies of Spirit’s size were bought whole, and the old servers — along with the email and all the archived documents — simply moved over to the new owner. This time, while other bits of the airline went to JetBlue, hedge funds, and other corporations, Google paid $10 million for part of Spirit’s dataset.

What was in the lot?
100 million emails across 80,000 mailboxes. 500 million messages in Microsoft Teams. 17.1 million files in OneDrive and 20.6 million in SharePoint. 30 million lines of code. 3.4 million payroll records. More than 175,000 personnel files dating back to 1986. The whole back-office picture: pricing models, booking curves, flight statistics, Wi-Fi receipts, and onboard purchases. Databases covering 7 billion competitor flights and 7.5 billion passenger transactions since 2008.
The passengers themselves aren’t included, at least not formally: the administrators will sell 97.5 million passenger profiles as a separate item from the bankruptcy estate. A specialized company will scrub names, addresses, and complaints from the employee correspondence before Google gets the archive — but Google itself selects and pays that company. The agent deciding what counts as personal and what gets deleted reports to the very party that will end up owning the data.
The Association of Flight Attendants-CWA has filed a technical, precisely targeted objection to the deal.
Anonymization, they argue, doesn't make the underlying event any less sensitive. What's more, the sale agreement explicitly requires preserving "referential integrity across the data set," the links that let you trace information through different systems. Without them, the archive loses its value — the buyer needs a working model of the organization, not a scattered pile of emails — but anonymization becomes a fiction if they stay intact. A scrubbed dataset will still reveal which crews racked up the most complaints, who flew together on recurrent training trips, and who got called in for a review. Spirit had just over 5,500 flight attendants. With a pool that small, the union says, reconstructing a dossier on a specific individual would be trivially easy.
The AFA-CWA wants everything relating to flight attendants pulled from the lot: personnel files, travel records, exam results, payroll ledgers, tax documents, crew rosters, and litigation materials. Its lawyers are demanding a strict vetting protocol and an outright ban on compiling digital dossiers on employees if the deal does go through.
That said, the union isn't challenging the sale itself. It isn't saying a dead company's data shouldn't be turned into merchandise, just asking to have its own people removed. Furthermore, no union is speaking up for the pilots, mechanics, dispatchers, ticket agents, baggage handlers, call-center staff, and everyone else whose emails sit in Spirit’s 80,000 mailboxes.
On September 9, Judge Sean Lane will decide whether or not the deal can go through. He’ll most likely approve it, just load it down with conditions, like pulling some disputed emails and banning digital dossiers. In the end, the creditors will collect their $10 million and Google will get the vast majority of the archive. The outcome itself really no longer matters, because the situation has already set a precedent. A company's digital legacy has gone up for sale as a separate lot and now has a legal status and a price tag.
Every year, thousands of companies go bankrupt around the world, leaving behind servers full of email — data that was never meant for outside eyes. While a company is alive, that archive is untouchable, but once it dies, the correspondence in the bankruptcy estate will most likely share the fate of Spirit’s.
The dragon no longer needs to fly anywhere to eat. All he has to do is file a bid on time.
Sources
References cited in this piece. Last verified on the published or revision date.
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09
- 10
- 11
- 12
- 13
- 14
- 15
- 16
- 17
- 18
- 19
- 20
- 21