Why Spirit Airlines’ internal data has become a hot commodity for AI companies

Spirit Airlines hasn’t flown since May, when the discount carrier announced an “orderly wind-down of operations” as part of bankruptcy proceedings. But the grounded airline might be set for a multimillion-dollar payout from a newly valuable asset: the sale of its corporate data, including internal wikis, emails, source code, and spreadsheets that can be used to train artificial intelligence.
Google successfully bid $10 million for the data trove, according to court documents, though AI training provider Micro1 says it has since submitted a rival $12.5 million offer. It’s unclear whether Micro1’s offer, which came after the formal auction had concluded, will be considered. An attorney handling the data sale in the bankruptcy didn’t immediately respond to an inquiry from Fast Company.
Under the Google deal, any personally identifiable information would be processed for privacy by a third-party “deidentification agent,” according to court records, and the dataset doesn’t include Spirit’s customer lists. Still, a Spirit flight attendant union has formally objected to the Google sale unless more steps are taken to protect employee privacy, arguing that workers can still be identified even after their names are removed from the files.
“We acquired part of an enterprise dataset from Spirit Airlines, which can be helpful in improving our products and AI models,” a Google spokesperson writes in an email to Fast Company. “We will not receive any personal information from this dataset.”
The Spirit bidding war is part of a broader rush by companies building and refining AI systems to acquire datasets from both shuttered and still-operating businesses. They say the records can help train AI to handle office work in realistic business environments.
“Companies are sitting on decades of records that show how real work gets done, and that data is now some of the most valuable material for training and evaluating AI,” writes a spokesperson for AI training company Mercor, which itself had offered $7.5 million for the Spirit data, in an email to Fast Company. “We partner with leading companies to license their operational data to the labs building the next generation of models. Spirit was that same process applied to a bankruptcy estate.”
Such corporate data transactions come alongside controversial pushes to acquire everything from Reddit’s forum archives to out-of-print books that could give AI models an edge, as well as hefty payouts to experts in fields from finance to poetry to help train and test AI. While AI companies say corporate records are invaluable for developing tools that can work in real-world business environments, privacy advocates caution that they could expose sensitive information about employees and customers even after steps are taken to anonymize the material.
“If any raw personal data goes into an AI training system, there’s always a risk that it could resurface in an output,” says Calli Schroeder, senior counsel at the Electronic Privacy Information Center and director of EPIC’s AI and Human Rights Program.
Corporations buying and selling information is nothing new, with data brokers dealing in records like customer mailing lists for more than a century. And, says Schroeder, data from bankrupt companies has previously been sold to businesses looking to better understand an industry or acquire new sales leads. (In some cases, regulators have pushed to limit the use of consumer data post-bankruptcy, citing privacy policies.)
But the growth of AI has created new demand for nuts-and-bolts operational records such as GitHub logs, Slack chats, emails, spreadsheets, knowledge bases, and customer support documentation that previously wasn’t particularly marketable.
“Particularly now that these big AI companies are just increasingly hungry for workplace data to train their AI models, there’s just a massive expansion of what counts as commodifiable data,” says Alexandra Mateescu, a researcher at the Labor Futures initiative at Data and Society.
Micro1 operates a program it calls Data Partnerships to acquire such corporate records. It uses the information to build realistic, simulated work environments for reinforcement learning, essentially honing AI’s ability to handle particular tasks and testing AI agents’ prowess, says founder and CEO Ali Ansari. AI agents in training are essentially set up as employees at fictitious companies generated from anonymized data from real businesses, equipped with realistic versions of real-world software tools and tasked with getting things done.
“It needs to have a world that it refers to, whether it’s company files, Slacks, documents, etc., so that it can do realistic actions and refer to realistic data,” Ansari says. “There’s a huge appetite from labs, from enterprises, from really all of our customers to gather lots of this data to build these realistic environments.”
Before the models in training touch the data, a separate AI-driven process strips out sensitive information, whether that’s email addresses and ID numbers or more complex personal data, he says. The raw, unredacted data is typically deleted within 30 days. Identifiers like email addresses are typically replaced with simulated dummy data in a similar format. Micro1 also tells data providers not to share various categories of sensitive information, ranging from trade secrets and legally privileged files to health data, he says, and excludes them from training sets when they do pop up.
“We try to actually exclude it proactively,” Ansari says.
In practice, any off-topic information in datasets, like, say, an employee talking about their boss on Slack, is often ignored by AI agents, he says, since it’s unrelated to the tasks they’re charged with completing.
In exchange for corporate data collections, the company typically pays anywhere from $100,000 to $2 million, though some datasets may drive even higher payouts, Ansari says. Micro1 also pays a $50,000 finder’s fee to anyone who successfully refers a new data partner, and the company typically hears from at least 200 companies per day interested in selling their data. It generally buys records from “tens of companies” in fields from film production to manufacturing “every week or two,” Ansari says. The company says it has committed roughly $20 million to enterprise data purchases, not counting the Spirit bid, in the past two weeks.
“In fact, just before this call, I approved $1.5 million in referral payouts,” Ansari told Fast Company in a Thursday interview.
Many Micro1 partners are small or medium-sized businesses with between 30 and 200 employees, meaning the data payout can be a fairly significant revenue source, Ansari says. Some companies also appreciate that they can contribute to the evolution of AI tools they’re already using, he says, as well as Micro1’s commitment to data privacy and security.
“We are in the business of ensuring that data is kept secure,” he says. “Data is kept very private, and that is really one of our core pillars.”
Still, privacy advocates caution that even anonymized data can often be linked to particular people through means ranging from the context of communications to particular speech patterns. In its legal filing, the Spirit flight attendant union argued that simply redacting names and other identifiers isn’t enough to safeguard worker privacy, pointing to a broad set of potentially sensitive records included in the Spirit dataset.
“A pseudonymized dataset can still disclose which crew bases generated grievances, how a small subset of flight attendants performed on recurrent training, which employees were subject to investigation, what compensation adjustments followed which events, and what employees said to one another about management, staffing, or their union,” the union’s legal finding argued. “The Assets Schedule includes free-form materials, including 100 million emails, 500 million Teams items, OneDrive and SharePoint repositories, litigation case files, and employment contracts, whose confidentiality inheres in their substance and context.”
An attorney for the flight attendant union didn’t immediately respond to an inquiry from Fast Company.
While Micro1 and the other companies interested in the Spirit data emphasize their commitment to privacy, experts say there are limited legal protections for consumers and, especially, employees concerned about whether their information or work product ends up in AI training sets. Recent reports have highlighted how AI systems can behave unpredictably in training, even bypassing safeguards meant to restrict their behavior or access to the internet. AI training providers also aren’t themselves immune to outside data breaches, with Mercor currently facing litigation after disclosing a March breach involving “sensitive information” related to experts it pays to train AI.
“It feels like a microcosm of a larger issue in this country, which is we don’t have strong, comprehensive privacy protections that protect people from this kind of stuff,” says Reem Suleiman, senior campaign director at Fight for the Future.
Some workers, like the business owners who contract with Micro1, may be excited about the prospect of helping improve AI tools they regularly use, and Ansari argues that AI will boost the number of jobs available to humans and generally improve working conditions.
“What I think is the most probable case is that pretty much all jobs will change for the better, and humans will have a great time working on the creative aspects of their job—the more fun parts—and AI will help with a lot of the rest,” he says.
But based on recent research surveying popular views of AI, some workers will likely be concerned that their work product will be used to train their future robotic replacements.
“A lot of this training for AI is basically being used to develop AI agents or chatbots to replace the very workers whose data has been commodified,” says Mateescu. “And those workers haven’t been compensated or asked for consent for that being used.”
After all, documents, chats, and emails produced at work typically belong to the employer, not the employees who create them. Another AI training company, Handshake AI, recently made headlines by offering to pay $6 per page for “real-world professional documents.” But a LinkedIn listing indicated contributors should only offer “documents that you own or are authorized to share,” and experts cautioned employees shouldn’t try to sell their employers’ documents without permission.
Still, some employment contracts, including union agreements, contain provisions related to protecting employee data. It’s possible that such provisions may become more prevalent, or that future laws or regulations could bring greater protections, though those seem unlikely to pass in the current political climate.
“I think there’s a mood right now in the government where you cannot regulate AI because it may hurt American companies’ competitiveness on the global market,” says Alice Marwick, director of research at Data & Society.
In the meantime, demand for AI training data seems likely to keep growing. Even as companies themselves adopt AI for more tasks, Ansari says, their operational data remains valuable because AI systems need to understand how workplace operations are changing in the real world.
“As we use the real data to improve models to then improve company operations by using these models, we then can buy more data from the new reality of how the companies operate,” he says.