NEWS

Google Buys Bankrupt Spirit Airlines' Data for $10 Million

Yellow Spirit Airlines jets parked in a desert storage lot, their tails marked spirit, with mountains in the background
Spirit Airlines jets in desert storage after the carrier ceased operations. Its data outlived the airline. Source: CNN
TLDR

What ten million dollars bought from a defunct airline

Spirit Airlines stopped flying in 2026 after years of losses, and in bankruptcy its data became an asset to be auctioned like gates or aircraft. Google won that auction, paying $10 million for a slice of the airline's internal records and, according to reporting from CNN, outbidding the data-labeling firm Mercor.

The dataset is unusually broad. It spans more than 100 million employee emails, hundreds of millions of Microsoft Teams messages, over 30 million lines of internal code, pricing data drawn from about 7 billion competitor flights, and roughly 7.5 billion passenger transaction records reaching back nearly two decades. It also carries operations, revenue, HR, marketing, audit, and fraud records, the day-to-day exhaust of running an airline.

We acquired part of an enterprise dataset from Spirit Airlines, which can be helpful in improving our products and AI models.
Google, on the Spirit Airlines data purchase

Google said the dataset will be "rigorously scrubbed of any personally identifiable information by a third party" before it receives the files, and that passenger profiles and loyalty program information were left out of the sale.

Dataset componentApproximate volume
Employee emails100 million and up
Microsoft Teams messagesHundreds of millions
Lines of internal code30 million and up
Competitor flights pricedAbout 7 billion
Passenger transaction recordsAbout 7.5 billion, nearly 20 years
Price paid by Google$10 million
What Google acquired from Spirit Airlines' bankruptcy estate. Source: CNN; Google statement.

Why a bankrupt company's inbox is now an AI asset

The number worth sitting with is not 7.5 billion. It is 10 million dollars. A frontier lab paid roughly the cost of a single senior researcher for two decades of how a real business actually priced tickets, negotiated, staffed, and audited itself. That is the kind of operational data the public web does not contain, and it is exactly what large models run short of once they have ingested everything freely available online. Bankruptcy is what makes it purchasable: a living company would never sell its email archive, but a dead one has creditors to repay and no reputation left to protect, which turns its most sensitive records into inventory.

The price also looks like a bargain next to what data now costs on the open market. OpenAI pays News Corp about $250 million over five years, and Reddit charges Google roughly $60 million a year for the same firehose it licenses to OpenAI for about $70 million. Against those recurring bills, a permanent copy of an entire airline's operational history for a one-time $10 million is one of the cheapest large datasets a major lab has bought.

Bar chart comparing disclosed AI training-data deal values: Reddit and OpenAI 70 million dollars a year, Reddit and Google 60 million dollars a year, News Corp and OpenAI 50 million dollars a year annualized, Dotdash and OpenAI 16 million dollars a year, Spirit and Google 10 million dollars one time, Financial Times and OpenAI 8 million dollars a year
Google's one-time $10 million for Spirit's data sits below what major publishers charge every year. Source: Reddit SEC filings, company disclosures, and reported deal terms.

The privacy promise deserves scrutiny. Stripping identifiers from structured pricing tables is straightforward. Doing the same across 100 million emails and hundreds of millions of chat messages, written by named people about named colleagues and customers, is far harder, and de-identification of free text has a long record of leaking the very details it is meant to remove. Google has moved the risk to a third party and a verb, "scrubbed," that is doing a great deal of work.

For everyone else, the signal is that proprietary data is becoming a moat you can now buy at a distressed-asset discount. As the open web dries up as a training source, the scarce input is real institutional behavior, and the companies with the deepest pockets are the ones positioned to acquire it when a competitor, or an unrelated business, collapses.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.