Why Smarter AI Models Could Drive Up Compute Prices 10x

The Compute Gap: Revenue Growth vs. Capacity Scaling

AI labs are experiencing a massive divergence between their revenue growth and their available compute capacity. For example, Anthropic's revenue has reportedly grown 10x year-over-year for three consecutive years, while total lab compute capacity only increases by approximately 3x annually.

To sustain this revenue growth while compute scaling lags, labs must rely on one or more of the following mechanisms:

  1. Increasing Margins: Expanding the profit margin on inference.
  2. Increasing Compute Prices: Raising the cost of renting or buying compute.
  3. Shifting Compute Allocation: Increasing the percentage of total compute dedicated to inference rather than training.

While all three are occurring—with inference margins for some models rising from 40% to 80% and OpenAI's inference spend increasing from 25% to potentially over 50%—labs are reluctant to prioritize inference. Labs view inference revenue as a means to fund the training of next-generation models; shifting too much compute to inference signals that AI progress has stalled and the lab has transitioned into a mere cloud provider.

Why Compute Prices are Likely to Rise

If labs cannot sustain 90%+ margins without being competed away, the primary "escape valve" for the revenue-compute gap is an increase in the price of compute. This trend is already visible in the market:

  • Spot Price Increases: Spot prices for compute are more than 40% higher than they were during the February trough of the current year.
  • Premium Pricing for Scale: Frontier labs require secure, large-scale clusters rather than spot instances. For instance, Google is reportedly paying $900 million per month for 110,000 GPUs (a mix of GB200s and GB300s), which is double the current spot price per hour.

The Monetization of Intelligence

As AI models become smarter, they can monetize the same amount of compute more effectively. If an H100-equivalent GPU could run a true human-level software engineer, the economic value generated by that hardware would far exceed current rental prices. Based on current software engineer salaries, such a GPU should theoretically rent for over $250,000 per year—more than 15x the current spot price.

While some argue that a massive influx of AI engineers would lower the marginal value of labor (the "lump of labor fallacy"), standard economics suggests that innovation and specialization typically keep the marginal value of high-skill labor high, implying that the marginal value of compute will remain high as capabilities increase.

Economic Implications of Higher Compute Costs

Rising compute prices and increased model efficiency create several strategic shifts in the AI ecosystem:

The Alchian-Allen Effect and Model Efficiency

According to the Alchian-Allen effect, when the cost of a scarce input (compute) increases, the incentive to economize that input grows. If compute is expensive, using a weaker, less efficient model is costly because it burns more tokens to achieve the same result. Consequently, labs that can train the most compute-efficient models can charge a significant premium because they are effectively "creating" more compute by reducing the waste of the scarce resource.

Market Concentration and Pricing Out Applications

Higher compute costs create a barrier to entry. Top labs can bid for compute resources more effectively than smaller competitors because they can extract more value from the same hardware. Additionally, low-value applications—such as generating "AI slop"—may be priced out of the market as frontier labs prioritize high-value automation, such as AI-driven AI research, which they are willing to pay more for.

Constraints on Compute Supply Scaling

The 3x annual growth in compute capacity is difficult to accelerate due to three primary bottlenecks:

  1. Moore's Law (1.4x): Hardware efficiency gains are slowing and may be difficult to sustain.
  2. Fabrication Capacity (1.2x): The construction of new fabs is bottlenecked by the availability of ASML EUV machines.
  3. Wafer Allocation (1.8x): AI is absorbing wafer capacity previously used for smartphones and PCs. At TSMC's leading-edge N3 nodes, AI allocation has risen from 60% to 86%, leaving little room for further growth via reallocation.

These constraints suggest that compute supply is inelastic and unable to absorb large demand shocks, making price increases more likely in the pre-singularity regime before robotic automation can fundamentally lower the cost of raw material processing and chip production.

Sources