Note · AI Implementation Strategies

Google's $10 million bid for Spirit Airlines' email:why I think the data is worth it

David He, FounderOctober 10, 20264 min read

Google bid $10 million for Spirit Airlines' old email to improve its AI models. Nate B Jones doubts it's worth it. I train AI agents daily, and I think it is.

Google bid $10 million for the emails, Teams messages and files Spirit Airlines left behind, and Nate B Jones doubts records like those are worth teaching AI agents with. I train AI agents to do tasks every day, and I think the data is worth the money.

If you've ever had an AI draft that sounded nothing like you, the missing piece was probably your own email. That's why I end up disagreeing with Nate on this one.

What Spirit is selling, and who objected

Spirit stopped flying on May 2. Its estate has been selling what's left, and in August it auctioned the data.

Google won on August 14 with a bid of $10 million, according to the notice of auction results filed with the bankruptcy court. Google told CNN the data can help improve its products and AI models.

The August filing lists:

  • about 100 million emails
  • 500 million Teams messages
  • 37,660,321 files: 17,082,644 in OneDrive and 20,577,677 in SharePoint
  • 667,563 IT tickets
  • 516 code repositories

Later filings narrow parts of that list. Passenger profiles were excluded.

Four days after the auction, the flight attendants' union objected. It says taking the names off a disciplinary letter or a leave request doesn't make the contents less private. It asked the judge to withhold approval unless those records are kept out or screened. The hearing was set for October 14, and on October 9 it moved to November 6.

Whether employees' records belong in the sale is for the judge, and I'm not arguing that here. My question is a different one: is the data worth what AI companies are paying for it?

Nate's case: the real work is in alignment

Nate B Jones covered the sale in an essay and a video on October 8. He used to write product requirements documents at Amazon, and now writes an AI newsletter.

His point is that much of what made the work matter never reached an email. "The real work is in alignment," he writes.

In the video he goes further. He expects agents trained on records like these to produce a simulation of work, and he doubts it's worth teaching them this way.

Normally I share Nate's view, but this time I disagree. It's worth the money.

Why I disagree: the emails tell my agent what to do

I train AI agents to do tasks every day. I show an agent how I already did the job, and I correct it when it's wrong.

Between June and October I approved 548 messages my agent drafted, and 62% went out without a single edit. For simple acknowledgments it was 86%. For money or contracts, 47%. I wrote up how I measured that in giving an AI agent authority to answer clients.

Then on October 6 and 7, five times, I told my agent to keep a reply I'd just approved as an example, and to send that kind without asking next time. Each example pairs a message with the reply I approved, and the agent uses it the next time that kind of message arrives.

Those emails are the exact thing that tells my agent what it should do.

Nate is right that email doesn't hold the full picture. The solution is to get the other forms of communication too. My agent also reads my texts and my call transcripts.

As proposed, this sale does that too. The Teams messages, files, tickets and code come with the email. The sale agreement also requires what it calls referential integrity: the records have to stay linked to each other.

One difference: I chose to hand my records to my agent. Spirit's employees didn't.

What I can't prove, and what it costs

My agent has me to correct it. Google gets the record and nobody to ask, and I haven't seen the data. Everything I know about what's in it comes from the court filings and the coverage of them.

I still think having part of the record is much better than having none when you're training an AI for a specific task.

At the bid price that's under two cents per message or file: $10 million across 637,660,321 emails, Teams messages and files is about $0.0157 each.

Two other AI data companies want it too. Mercor was named the alternate bidder at $7.5 million. Another, micro1, told the court it would pay $12.5 million.

Anyway.

An email is a partial record of a job. It's still the first thing I'd hand an agent that had to learn that job.

If you want to try it

You don't need an airline's archive to do this. Here's what I do:

  • Pick one task you do.
  • Give the agent the emails where you already did it.
  • Add your texts and call transcripts on the same work, because email misses part of it.
  • Correct the agent when it's wrong, and keep the replies you approve as examples.

If you had to teach an agent one task you do, which of your own records would you hand it first?

More notesnewest first

Working on something like this?

Bring the app or the process to a free 15-minute call. I will tell you what I would look at first, and whether I am the right person for it.