Colossus: why Anthropic will probably keep paying $15bn a year

Compute
Finance
Author

David Leitch

Published

August 10, 2026

Examining the bigger picture around the SpaceX Colossus lease

In brief

  • AI is useful and very rapidly becoming more useful. Initially mainly used for computer coding, it can now perform many knowledge management tasks and will likely continue to improve. I speak from experience: I have been using AI since pretty much the first commercial models became available, so I have watched the capabilities increase.
  • But AI is not yet easy to use well. It’s easy to spend money and come up with a good-looking diagram; getting the most out of it takes real effort. As capabilities increase further it’s likely to become easier to use.
  • People will pay to use AI, and inference — the business of actually helping them — is already profitable, in some cases very profitable. So investment will grow.
  • Improvement in model capability requires more brute compute power for training runs, and that means more MW. It’s plausible that continuing the current rate of improvement will require two to three times as many training MW in two years’ time as today. It will also require continued and hard to forecast improvements in algorithmic methods.
  • As a result of all of the above I think it’s more likely than not that Anthropic will wish to keep paying the US$15 bn a year it pays SpaceX for its Colossus data centre capacity. It won’t, in my opinion, need the capacity for inference, but it will most likely need it for training.
  • This note does not go into the environmental costs of the USA rushing headlong into a gas fuelled data centre explosion. In the USA over 2025 and 2026 something like 37 GW of data centres have been announced. My estimate is that at least 11 GW of that cohort is under construction or already complete. That’s the next research item. On the power side rather than the data centre side, Cleanview reports about 90 GW of behind the meter generation announced in the USA (40 GW in TX) of which 2 GW is operating, and looks to another 10 GW in 2027. Like I say, it’s another piece of the jigsaw.

Inference is profitable

One of the prime sources I follow is suggesting that inference on its own is fairly profitable. I looked at this by considering the Anthropic/SpaceX Colossus data centre rental.

Colossus is roughly 325,000 GPUs across SpaceX’s Memphis and Southaven sites — about 540 MW of compute draw — leased on a monthly fee that runs to May 2029 but is cancellable by Anthropic or SpaceX on 90 days’ notice.

The difficulty in all these numbers is the lack of disclosure. For anyone who is not an insider, the analysis is built end to end on what is at best informed speculation. Specifically, we have press reports of June quarter Anthropic revenue. We don’t know the capacity split between inference and training, and so it’s hard to work out how much inference is done on Colossus GPUs.

To work out how much could be done you need three things: average tokens per $ of revenue, how many tokens a GPU can serve, and what utilisation those GPUs run at. The last is much the hardest, and the only inference — boom boom — comes from the off-peak batch discount that Anthropic and OpenAI both offer. Neither would discount 50% for off-peak work unless that compute were otherwise sitting idle.

With all that stipulated, I put Anthropic’s inference gross margin, after Colossus rental costs, at maybe 35% — or US$8 bn on a full-year basis. The table below is on a GPU basis.

Figure 1: Anthropic inference economics against the Colossus lease. Source: ITK estimate

Whether Anthropic keeps paying depends on two things: whether SpaceX can put the capacity to better use, and whether Anthropic’s own demand keeps growing faster than its supply.

A second way of looking at the deal is as an option. It cuts both ways: Anthropic can walk on 90 days’ notice, and SpaceX can decline to renew. The second half of Figure 1 tries to bound what it would take for SpaceX to push Anthropic out.

There are only two plausible routes: a version of Grok good enough to take market share at top-model prices, or incentives that push Cursor users onto Grok rather than Anthropic or OpenAI models. Cursor is worth a lot of money on the price Musk paid, presumably it offers value beyond going to a model provider directly. My own feeling is that Vibe coding will continue to take coding market share.

SpaceX will struggle to get people to abandon their favourite models, because I don’t think Musk’s reputation carries the trust required. Model choice is about more than cost against ability — being cheaper won’t make people who value frontier capability switch, and there are plenty of cheap models already.

The price paid for compute relates to availability and duration and therefore lessor risk

SpaceX bears the risk: it is writing a short-term lease on fully equipped data centres that are available now. But if Anthropic cancels, SpaceX is exposed.

SpaceX is building a hedge against cancellation risk by developing a more competitive version of Grok or alternatively by moving the Cursor volume onto the Colossus stack. If I could see the value in Cursor more clearly that might be the more credible defence.

Figure 2: Disclosed rates for AI data centre capacity against the period the counterparty is contractually bound. Source: SEC filings and operator disclosure

Can Grok close the gap?

It may be that future versions of Grok will be much better, although I myself doubt that Grok can easily win trust. But so far, relative to the competition, Grok has been falling further behind, not catching up.

Figure 3: Grok against the frontier at each release. Source: Artificial Analysis via ITK model index

Interpreting that gap remains a matter of judgement. It’s a points measure on the Artificial Analysis Intelligence Index. The index methodology is revised over time; the most recent re-grade lifted the frontier models by about two points while leaving the tail largely unchanged.

However this is only one version of the leaderboard. Different leaderboards give different views.

Figure 4: Three scoreboards, one week: where they agree and where they do not. Source: Artificial Analysis; Epoch AI; LiveBench

And no leaderboard says what a model costs to run, which at the frontier separates the models more than capability does.

Figure 5: Model capability against the cost of running it. Source: Artificial Analysis

Why capability improvement keeps demand tight

Even as progress remains powerful, the leading model changes.

Figure 6: The frontier moved from 20 to 60 in seventeen months. Source: Artificial Analysis Intelligence Index

My current understanding is that it takes a ten-fold increase in compute to add roughly 33 points to the frontier score, and that each further 10 to 14 points roughly doubles the length of job a model can complete.

Figure 7: Ten times the compute, ten times the job. Source: Artificial Analysis Intelligence Index; METR Time Horizon v1.1; Epoch AI

So the fact that model improvement now runs on electricity supply growth is why I think data centre demand stays tight.

Figure 8: Anthropic compute capacity to end-2027, and what share of it the Colossus lease is. Source: ITK estimate

On my estimates Colossus falls from 28% of Anthropic’s fleet today to under 10% by end-2027. That looks like a lease you could drop — until you notice that a single frontier training run is heading for roughly the same 540 MW. Anthropic will not be short of megawatts for inference but it may need all its capacity for the increasingly large trains runs and thats why it may continue to pay US$15 bn a year.