Alex Karp isn’t crazy, and he isn’t just talking his book. On July 1, the Palantir CEO went on CNBC and asked the right questions about model providers: who owns the data, where is it cached, is anything transferred back to the provider. Too many people focused on his style and missed his point – they are stealing your alpha.

OpenAI and Anthropic tell commercial customers they will not train on their data. After I recently moved my company off Anthropic, following a Supply-Chain Risk designation, I read the actual agreements and found a hole big enough to drive the entire AI industry through.

The OpenAI Services Agreement says: “OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.” Reasonable, until you check the definitions. “Customer Content” means the Input and the Output. “Input” is what the customer sends the model. “Output” is what comes back based on the Input. Here’s why that’s vaguer than it sounds.

The hidden tokens

Early large language models generated answers token by token without scaling their effort to the difficulty of the question. Humans don’t work that way. Ask someone what 2 + 2 is and they’ll answer instantly. Ask them to plan a family reunion, and they’ll think it over first.

Newer reasoning models do the same thing, spending computation on intermediate steps before returning a final answer, breaking a hard question into parts and working through each one. That intermediate reasoning produces real data on the way to the final output, which raises the question: is it actually capital-O Output, legally? Nothing in these agreements says so.

Who owns what

The definitions matter because another clause ties ownership directly to them. OpenAI’s agreement states that the customer “retains all ownership rights in Input” and “owns all Output,” with OpenAI assigning its interest in Output to the customer.

Picture handing a consultant your confidential financial forecast and asking for next year’s headcount budget. The consultant reasons by filling a notebook with calculations, then hands you a one-sentence answer, and keeps the notebook. Your contract with the consultant covers the question and the answer. It says nothing about the notebook.

That’s what happens every time you use a reasoning model. It generates intermediate reasoning tokens before returning a final answer: a digital scratchpad of facts pulled from your prompt, intermediate conclusions, and newly derived insights about your business. The labs know it’s valuable. OpenAI has said it hides raw chains of thought for reasons that include safety, user experience, and competitive advantage. Anthropic bills customers for full internal thinking even when none of it is ever shown to them.

So you pay for the production of this data. You never see it. And no one will say whether it’s legally yours. OpenAI’s no-training promise covers Input and Output, but the notebook fits neither definition. If reasoning tokens are Output, say so, and extend the protections to cover them. If they’re not, OpenAI has created a third category of data its agreement never defines and never protects.

Anthropic has the same problem. Its API bills for reasoning tokens that never appear in the visible response, and returns the raw reasoning encrypted, so only Anthropic can decode it. So do you own it? This isn’t a drafting oversight. It’s a convenient ambiguity, given how much this data is worth.

Distillation proves the notebook is valuable

The labs can’t dismiss reasoning tokens as meaningless computational exhaust. Their own conduct proves otherwise. Anthropic says DeepSeek, Moonshot, and MiniMax generated more than 16 million Claude exchanges to help train competing models, calling it industrial-scale distillation, a shortcut to capabilities that would otherwise take enormous time and money to build. Labs protect outputs aggressively because outputs transfer intelligence, and raw reasoning tokens are an even richer record of how a model reaches an answer. When OpenAI launched o1, it said this directly:

“After weighing multiple factors including… competitive advantage… we have decided not to show the raw chains of thought to users.”

They want it both ways: bill you for the reasoning, hide it because it’s strategically valuable, and decline to say whether it’s legally your Output.

The fair use playbook

The double standard is hard to miss. This industry was built on the argument that a company can ingest someone else’s protected work, transform it through training, and own the resulting asset. Apply that logic to your enterprise data. You provide confidential material as Input. The model transforms it into reasoning tokens that aren’t identical to your Input but are valuable, newly generated material derived from it. Why would anyone expect the labs to resolve that gray area against their own interests?

As Karp put it, “the jig is up.” You don’t own your alpha. The guy everyone called erratic was trying to tell you. Even Zero Data Retention doesn’t close the gap: both OpenAI and Anthropic offer it, but it requires separate approval and doesn’t guarantee reasoning data gets discarded rather than retained. If not keeping your data takes a special request, retention is the default for a reason.

Which is it?

I expect intermediate reasoning tokens to be assigned to me, the same way I’d want my own notebook back. I may need to reason from it again. The labs haven’t clearly assigned those rights, because doing so means giving up the value. And if reasoning tokens aren’t Customer Content, it’s entirely possible they’re being used to train the next model. Nothing in the labs’ conduct gives me confidence otherwise.

Sam Altman and Dario Amodei can resolve this in a sentence each. Until they do, Karp was right. You own the prompt. You own the answer. They keep the notebook.

Adam Fish is the CEO and co-founder of Ditto, an edge data platform built for unstoppable operations. Its peer-to-peer sync technology keeps applications running with no cloud dependency, in U.S. defense programs including Special Operations Command and for commercial brands like Chick-fil-A.

The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.

This story was originally featured on Fortune.com

Read More