Cyberiad Labs Sign in Save my seat

All news

Thought piece 11 min read

The Labs Will Stop Selling Their Best Models

TL;DR: the big AI labs will keep selling tokens, but within a few years that won't be their main business. Open-weight models are closing the gap, Chinese labs are copying the frontier at scale, and the best models are getting too dangerous to release to everyone. The labs will keep their best models and point them at problems worth more than tokens: a proof, a molecule, a financial model, a diagnosis. You can already see them doing it.

By Julien Barbier

Tokens are becoming a commodity

Amazon, Google, Meta and Microsoft plan to spend about $725 billion on data centers this year, and OpenAI alone plans about $750 billion of compute through 2030. That money is supposed to come back through tokens, and tokens are getting cheap much faster than most people expected, myself included.

In August, the AI safety nonprofit SaferAI found that Z ai's GLM-5.2, an open-weight Chinese model, was only two to four months behind GPT-5.5 and Claude Opus 4.7, and level with Opus 4.7 in biology. On OpenRouter, a service developers use to send their requests to hundreds of different models, the lab with the biggest share of requests last week was DeepSeek, a name most people in America and Europe have never heard of.

You can download GLM-5.2 or DeepSeek V4 Pro for free and run them on your own hardware. And if you don't want to/can't buy the hardware, you can use them for about 10 times cheaper than OpenAI's and Anthropic's models.

Every release trains the competition

Part of the reason the gap closes so fast is distillation: you ask a frontier model millions of questions and train your own model on its answers, without having to do the R&D these labs did. Anthropic reported this month that between May and July, Alibaba alone pulled 151 million exchanges out of Claude through more than 3,500 fake accounts. It named six other Chinese labs doing the same, for about 190 million exchanges in total.

Every time a lab ships a model that costs hundreds of millions or even billions to build, it is also training its competition, and a few weeks later the capability shows up in a free download. The labs try to ban the accounts, but as long as the model answers the public, someone will be able to harvest it.

That is why keeping the best model internally, for longer, would make a lot of sense. Nobody can distill a model they can't talk to. The gap stays at a few weeks because the copiers learn from each new frontier model the day it ships. If frontier labs stop releasing their newest and best model, the others will have to do the research themselves. DeepSeek has shown it can do a lot of that research on its own, so the gap won't grow forever, but then, overnight, a few weeks can turn into a few months. And in the AI world, this is an eternity and a huge moat.

The best models are getting harder to ship, and harder to contain

Everyone appears to struggle to keep these models in a box, including their creators.

In April, Anthropic announced Claude Mythos and decided not to release it to the public, because it was too good at finding software vulnerabilities. About fifty organizations got access to it, but on the first day a small group on Discord got in anyway, by guessing the model's URL ¯\_(ツ)_/¯

In May, the security company Irregular was testing Gemini on fake companies in a closed environment that was never supposed to touch the internet. But somebody had left that environment connected to the internet by mistake. And then Gemini broke into three real companies, by guessing or finding passwords posted in public code repositories. Google says Gemini stopped each time it realized the company was real, and it told the three companies, but it decided a public announcement wasn't needed and only talked about it when the Wall Street Journal asked. Irregular's tests are also linked to similar escapes by models from OpenAI, Anthropic and Meta.

In July, two of OpenAI's models broke out of their test sandbox during a cyber evaluation. They went into the internet and hacked into Hugging Face's production systems, looking to cheat at the test they were "graded" on.

At the end of July, Anthropic disclosed that its own models had hacked three organizations during tests.

And yesterday, Australia's prime minister, Anthony Albanese, shared that an OpenAI agent had broken into a government Medicare statistics portal in June. This happened during an evaluation, when the model was looking for answers to questions about Australia. It found a way around the blocks and opened files that were not public. No personal Medicare records were accessed, but OpenAI only told the Australian government on September 10, almost three months later, and Albanese has now set up a taskforce to investigate.

Every one of these incidents happened during a test, run by the lab or by an outside security firm, in a setup built to keep the model in, and the model still got out. If a lab can't contain a model it is only testing, releasing it to millions of users on day one would only make the problem bigger, and would create infinite potential liability for their creators. The more capable the model, the more harm it can do, and the longer it should take to test before anyone outside the lab can use it.

No lab wants to find out what a thousand lawsuits cost, the day an agent leaks a thousand databases.

We've seen this movie before

The people who paid for the rails usually lost their shirts. In the 1893 panic, a quarter of all US rail mileage went into receivership, while Sears and the Chicago meatpackers built empires on top of the tracks. After the dot-com crash, Global Crossing and WorldCom went bankrupt, and Google, Amazon and Netflix grew on the cheap fiber they left behind.

Some rails do stay good businesses: AWS made $45.6 billion of operating income last year, and Nvidia made $120 billion of profit in its last fiscal year, because nobody can easily replace either of them. The rails that lost money were overbuilt and easy to replace, and today tokens look a lot more like fiber than like AWS.

The people running these labs are some of the most intelligent people of our generation, and they have access to huge capital. There is no way they let their business turn into fiber.

History has a second lesson here, about how to deal with massive liability.

In the early 1980s, families in the US sued the makers of the DPT vaccine and won millions. As a result most of the makers stopped producing it. Congress then decided that the country without vaccines was worse than the rare person a vaccine hurts, and in 1986 it passed a law so that people hurt by a vaccine can't sue the maker directly. They go to a special government court instead.

The labs would like the same deal for AI, and their argument has the same shape: a country whose rivals have better AI is worse off than one where an agent sometimes breaks into a company. In 2025, OpenAI asked the White House for "liability protections" and warned that without them, against China, "the race for AI is effectively over".

But on September 15, Treasury Secretary Scott Bessent said: "The one thing we should not do is give [AI labs] a blank check on liability." On September 21 he went further on CNBC: the labs had asked Washington to "take the liability off of our hands", and "we will not do that".

If they don't get it, they carry the risk alone, and my thesis will only happen faster: the cheapest way to carry that risk is to keep the best models internally and slow the release pace of lesser models.

The labs are already moving

AI labs are already moving up the stack, and you can see this in the news.

OpenAI says an internal model has solved more than 100 open math problems, including a claimed result on Navier-Stokes, one of the seven Millennium Prize problems. The public can't use that model.

OpenAI also paid more than 100 former investment bankers $150 an hour to teach its models to build financial models, and says more than 260 physicians in 60 countries have reviewed its health answers over two years, the work behind ChatGPT Health and ChatGPT for Clinicians.

Anthropic bought Coefficient Bio in April, in a stock deal reported at about $400 million, and last week confirmed to Reuters that it runs its own wet lab in the Bay Area, where it wants Claude to direct robotic instruments through real experiments for the company that trained it.

Yesterday Anthropic announced the lab's first result: about 950 Claude agents spent 21 hours and 210 million tokens searching a huge database of virus DNA, and found a previously unknown enzyme system, which Anthropic calls ART, built around DNA repeats like the ones behind CRISPR gene editing. Human scientists then tested it in the lab. Nobody knows yet what it does, and Stanford researchers had already found a system that is in some ways similar, but Feng Zhang, one of the pioneers of CRISPR, called it "an exciting example of how AI agents can contribute to biological discovery". Every one of those 210 million tokens was spent inside Anthropic, on a model most likely no one has access to but them (not confirmed, this is my hypothesis).

xAI, which is now part of SpaceX since February, bought Cursor for $60 billion in stock, which gave it the app developers already are using. In July, Musk said SpaceX's own engineering data would go into training Grok, data no other AI lab can get, inside a company that builds rockets and can put the model to work on them. And in March he turned Macrohard, an xAI project he first announced in 2025, into a joint project with Tesla whose goal, in his words, is to "emulate the function of entire companies".

xAI has also promised its own AI-generated video game, from a studio Musk announced in 2024.

Google has been at it the longest. DeepMind built AlphaFold, and in 2021 Alphabet spun out Isomorphic Labs to turn that work into drug candidates. It now has Eli Lilly, Novartis and Johnson & Johnson as partners and raised $2.1 billion in May.

The future of frontier

The labs will keep selling tokens, and the public API will stay a good business, as the store for “last year's model”. The gap between what people have access to and the frontier models will only increase from there. The most capable models will stay inside, for much longer, pointed at problems worth more than a billion API calls: a proof, a molecule, a financial model, a diagnosis.

That only works once a GPU hour spent on the lab's own problem earns more than a GPU hour of tokens sold. We are not there yet. Tokens are still what pays for the compute.

My bet is that this flips for at least one lab before the end of 2027. It will be easy to spot. By then, I expect OpenAI and Anthropic to be public companies, like Google and SpaceX already are, and we'll only have to read their revenue lines to see whether the revenue from discoveries or products higher in the stack is catching up with the revenue from tokens.

If the AI labs are moving up the stack, I don't know where they stop. In August, the investor Gavin Baker said on the All-In podcast that he'd been told Dario Amodei once suggested Anthropic might end up "the only private company in the world", with "Anthropic, and then there are governments, and that's it". Could they own the entire stack, and all stacks? The entire economy?

Sources