Tokens Too Cheap to Meter – Why AI Compute Costs Are Plummeting

Token Costs Are Dropping Faster Than Any Past Trend

The price of completing an AI task has dropped about 2.5 orders of magnitude in the last 12 months, driven by simultaneous gains in hardware efficiency, model architectures, and inference engines. This makes the cost of a token so low that the overhead of invoking external tools (e.g., grep, HTML parsing, or a cargo build) will soon dominate the total expense of a workflow.


Hardware Efficiency Gains

GPUs double efficiency every two years

A logarithmic plot of GPU energy efficiency (GFLOP/J) shows a 1.3‑log slope, meaning efficiency doubles roughly every two years—comparable to the historic pace of Moore’s Law. The trend holds across generations, from early consumer GPUs to the latest data‑center H100.

"This is an increase in efficiency that we haven't seen since Moore's Law in the 1960s." – post author

Inference engines improve 10‑50 % YoY

Open‑source engines such as vLLM have achieved a ~40 % reduction in Joules/token between version 0.5.4 (Sept 2024) and 0.11.1 (Dec 2025). NVIDIA’s full‑stack MLPerf stack reports up to 50 % efficiency gains from version 2.0 to 2.1, while Intel’s MLPerf 6.1 shows a 2.4× throughput increase over 6.0.


Model‑Level Cost Reductions

Per‑task cost falls while per‑token price stalls

Frontier models still charge similar per‑token rates, but larger models finish tasks with far fewer tokens. The Pareto frontier of quality vs. cost (2025‑2026) demonstrates that the same benchmark can be run 100× cheaper today than a year ago.

Mixture‑of‑Experts (MoE) cuts compute

MoE architectures deactivate expert layers when unnecessary, yielding up to 7× fewer parameters for equivalent benchmark performance. Although local deployment still requires all experts in memory, the overall quality‑per‑joule improves dramatically.


Specialized Models Accelerate the Trend

Jev and Laya achieve "free" output tokens

TypeSafe’s Jev classifier charges $0.042 / MTok for input tokens and $0 for outputs—about $42 per billion tokens, roughly the cost of a few cents to read five books. Open‑weight Laya matches or exceeds Jev’s accuracy when fine‑tuned, though it requires more setup.

"We can’t prove it isn’t subsidized; we’ll need the long‑term to prove the sustainability of our pricing (which we expect to go down, not up)." – TypeSafe pricing note


When Tokens Become Cheaper Than Tool Calls

A back‑of‑the‑envelope calculation using GPT‑5.6 Luna ($0.30 / M tokens) and a typical MacBook Air (10 W idle, 30 W heavy) shows:

Tool Power (W) Duration (s) Cost (¢) Orders of magnitude cheaper than a Luna turn
grep 10 0.1 0.000007 4.5
HTML parse 10 1 0.00007 3.5
cargo build 30 30 0.00625 1.5

At current rates, a LLM call will soon cost less than a simple grep, making it economically attractive to embed models inside traditional tooling pipelines.


Economic Counter‑forces

Supply‑side Jevons paradox

Higher efficiency fuels induced demand: AI providers invest in more compute because each dollar spent yields higher revenue. This mirrors the classic Jevons paradox where cheaper electricity spurred greater overall consumption.

Business‑model viability

Critics note that cheap tokens alone do not guarantee profitability. Providers still rely on frontier‑model premium pricing, enterprise contracts, and volume‑based services. Open‑weight competition could pressure margins, but hardware vendors (e.g., NVIDIA) and specialized inference services will continue to profit from scale.


Future Scenarios

  1. Ubiquitous local inference – Within 3‑6 years, frontier‑quality models are expected to run on commodity hardware thanks to RAM‑efficient architectures (Mamba, quantization to 4‑bit) and MoE scaling.
  2. AI‑augmented tooling – As token cost falls below tool‑call cost, developers will replace ad‑hoc scripts with model‑driven assistants (e.g., jgrep). Build schedulers, security scanners, and CI pipelines may become AI‑first.
  3. Shift in software value – The primary differentiator will move from raw functionality to quality, security, and integration. Legacy SaaS products with high switching costs will retain advantage, while malleable, AI‑generated software proliferates in niche domains.

Community Reactions

"Tokens become cheaper than tool calls… I think this is a good time to invoke Stein's Law: 'If something cannot go on forever, it will stop.'" – @jetrink

"The Pareto chart is meaningless without a value judgement; the 'most attractive quadrant' misleads readers." – @meatmanek

"Energy efficiency graphs can be fitted with many trend lines; the claim that GPUs are doubling every two years needs more rigorous statistical backing." – @empw

"If inference continues to become a commodity, hardware margins will shrink, and profit will shift to services and specialized chips." – @bryanlarsen

These comments highlight skepticism about perpetual exponential trends, the need for careful interpretation of benchmark visualizations, and concerns about long‑term business sustainability.


Takeaway

The convergence of hardware efficiency, model architecture advances, and software engine optimizations has driven AI token costs down by over two orders of magnitude in a single year. This makes the cost of a token cheaper than many traditional compute primitives, foreshadowing a future where AI is embedded directly into the fabric of software tooling and infrastructure.

Sources

Related