
Google has agreed to pay $10 million to acquire a massive archive of data belonging to the now-defunct Spirit Airlines. The sale, which is scheduled for a court hearing on Wednesday morning, was first reported by Reuters on Monday. Every public account of the deal rests on a single word: deidentified. The data, according to the buyer and seller, has been scrubbed of personal identifiers, so nobody need worry.
But the data sale agreement filed with the bankruptcy court on 14 August reveals three crucial details about that deidentification process that have gone unreported. Read together, they change what the word is worth.
Google picks the firm that does the scrubbing
Coverage of the sale has described an independent third party cleaning the data before Google receives it. The actual contract is far more specific. Spirit must deliver the material to one or more third parties that are either acceptable to or designated by the buyer. In plain language, Google chooses the agent responsible for anonymising the data.
Google also pays for the privilege. The agreement makes the buyer solely responsible for every cost associated with deidentification, and states that those costs do not reduce the purchase price. The $10 million headline figure is not the final bill.
Google then gets to inspect the work. The contract stipulates that Spirit must give the buyer a reasonable opportunity to review and comment on the anonymisation process, and that the seller must give good faith consideration to those comments. The ultimate standard is a scrub that is reasonably satisfactory to Google.
None of this is inherently improper. It is standard commercial drafting in corporate asset sales. But it is simply not what an independent audit sounds like to the average person.
The scrub has to keep the records linked
The clause that matters most has been quoted by nobody. It concerns what the anonymisation agent must certify. The agent must issue a certification that the work meets the California Consumer Privacy Act standard. Health-related material must also conform to the federal health privacy rule. Both certifications apply whether or not those laws would otherwise reach the data, which is a genuine protection for individuals.
Then the sentence ends with a condition. The certification must hold “while preserving referential integrity across the data set”.
Referential integrity is a database concept meaning that the relationships between records survive the anonymisation process. The joins survive by design. One pseudonymous person still runs from an email address to a support ticket, to a code commit, to a payroll record.
That is exactly what makes the archive valuable for training AI agents. An agent that can follow a customer complaint through the entire operational chain is far more useful than one trained on disconnected text snippets. But referential integrity is also the property that makes any anonymisation fragile. If a pseudonymous identity can be traced across multiple data types, the risk of re-identification rises dramatically. The contract explicitly requires that fragile property.
What is actually in the box
The schedule of Spirit Airlines data is far more specific than the summaries that have appeared in the press. It lists 100 million emails across 80,000 accounts and 500 million Teams messages. Add 17,082,644 OneDrive files, 20,577,677 SharePoint files and 667,563 IT tickets. James Nani of Bloomberg Law first reported the headline volumes.
The engineering side runs to 516 repositories and roughly 30 million lines of code. The archive also contains 372,585 commits, 43,170 pull requests, and the pipeline logs surrounding them. This is a complete snapshot of a digital airline.
The operational data is enormous. It covers 763,391 flights and 5,014,676 crew pairings. It holds 190,312,864 booking records and 7,510,221,520 transactions reaching back to May 2008. Disruption and reaccommodation add another 3,000,347,472 rows. The volume is staggering.
Then there is the corporate interior: board presentations, budget walkthroughs, deal pipelines, due diligence reports, investment committee papers, lender materials and merger fairness opinions.
And the staff
The schedule also lists 175,658 employee records, with the system of record running from August 1986. It adds 3,426,618 payroll records and 148,018 employee tax forms. Then come 1,092,000 time cards, training records, recruiting files and travel requests.
Roughly 17,000 people lost their jobs when Spirit Airlines stopped flying on 2 May. Their correspondence, pay history and tax paperwork are now line items in a bankruptcy schedule. They signed employment contracts, not data licences. In a Chapter 11 estate, the legal distinction between an employee contract and a data licence does not arise.
What Spirit kept, and what it may still sell
The schedule marks the customer side as not included throughout. The excluded volumes are large: 97.5 million customer profiles, 50.2 million Free Spirit members, 740,000 card holders, 30,865,471 call recordings and 15,784,473 chat sessions.
Regulatory records sit on the same side of the line. That includes 2,491,715 disability service requests, denied boarding data and complaints to the Department of Transportation.
One clause in the agreement is worth reading closely. Spirit may not sell the assets to anyone but Google, with a single exception: it may sell its customer data list, including individual traveller spend aggregated by year, to buyers in the hospitality or travel industries.
So this sale did not shield the passengers. It separated them. The estate kept the right to market their data elsewhere. Axios reported the exclusions on Monday.
The bidding, and what it revealed
The auction for the data was contested. Google opened at $5 million. Mercor, an AI data company, countered at $5.2 million, then offered $7 million if it could take the raw data first and anonymise it itself, according to Business Insider. Google closed at $10 million.
The bidding pattern is revealing. A bidder priced the unscrubbed version above its own bid for the scrubbed one. Mercor, which sought a $20 billion valuation in July, remains the backup buyer at $7.5 million.
Google had been in diligence for a while. The agreement references a confidentiality agreement with Spirit dated 18 June.
Why an airline
The value of the archive is not text to pretrain on. It is a chain of consequence. The archive connects a ticket to the emails about it, the commit that followed, the review thread and the operational result. Agents completing multi-step work need exactly that kind of relational data, and scraped text from the public web cannot supply it.
Google has airline-specific reasons too. Google Cloud signed a five-year partnership with Ryanair on 12 August covering fleet operations and maintenance scheduling, suggesting the company is building toward aviation-focused AI products.
The purchase also fits a broader pattern. Google is in talks to pay $1.5 billion for a 35-person startup building coding environments, and China hit the same wall on training material this month. Meta tried collecting this kind of material from live staff and paused the programme after an internal revolt. An estate has nobody left to object.
What would settle it
Almost nobody can object now. The deadline for written objections passed at 4pm on 17 August, and the court required anyone attending to register by 11am on 18 August.
Judge Sean H. Lane hears the Spirit Airlines data sale at 11am on 19 August, over Zoom. Two questions outlast the hearing. Who audits a deidentification that the buyer designed, paid for and approved? And who buys the customer list, given that the estate kept the right to sell it separately?
Source:TNW | Privacy News
