How Anthropic used Claude to make claude.ai 3x faster

Anthropic reduced the load times and interaction latency of claude.ai and the Claude desktop app by approximately 3x during a two-week sprint. By deploying an internal research model (comparable to Opus 5.5) via "Claude Tag" into a dedicated Slack channel, the team enabled the AI to autonomously identify bottlenecks, create deterministic benchmarks, and ship optimizations with human oversight.

Performance Gains and Core Metrics

Anthropic focused on four primary user journeys that account for 95% of user activity: launching the app, starting a conversation, loading an existing conversation, and sending a message.

At the 75th percentile, the team achieved the following improvements:

  • Fresh load of claude.ai: Time to a typeable page decreased from 3.1 seconds to 0.55 seconds.
  • Claude Code session start: Latency dropped from 0.8 seconds to 0.3 seconds.
  • Claude Cowork cloud session load: Loading time decreased from 2.6 seconds to 0.73 seconds.

The "Measure to Optimize" Loop

The core technical insight of the sprint was that providing the AI with a precise measurement makes a performance problem tractable. Anthropic established a loop where Claude acted as the primary driver of optimization:

  1. Identification: A human would flag a slow experience (e.g., via screen recording).
  2. Benchmarking: Claude would trace the flow and build a benchmark to quantify the problem.
  3. Implementation: Claude would submit PRs (often several, sized for risk) to improve the metric.
  4. Verification: Claude would monitor the deploy and read field data to verify the win.
  5. Ratcheting: Once a win was confirmed, Claude would "ratchet" the benchmark down in CI, meaning any future PR that increased that specific metric would fail the build.

Moving Beyond Wall-Clock Time

Because wall-clock time is noisy and flaky for CI gates, the team shifted toward deterministic measurements. Claude implemented several "lab" benchmarks, including:

  • JS Instruction Counts: Using Valgrind with node --predictable to get exact instruction counts for pure-JS hot paths.
  • Browser Metrics: Tracking React commits per interaction, V8 function call counts, layout/style-recalc counts, and DOM mutations.

For example, Claude identified that a quarter of the instructions in the conversation message tree assembly were megamorphic dictionary lookups. By optimizing these, Claude reduced instructions by 48% and wall-clock time by 78% on that specific path.

Key Technical Optimizations

To achieve the 3x speedup, the team implemented a variety of architectural and micro-optimizations:

  • Static Composer: To eliminate the wait for React initialization, Anthropic baked a static HTML composer into the page. Users can type immediately while React paints over the static version. To prevent layout shifts, Claude built a test suite that asserts alignment within 1 pixel across 14 viewport sizes.
  • V8 Code Cache: Precompiled a V8 code cache for the desktop shell to prevent the main process from recompiling from scratch.
  • Render Optimization: Reduced sidebar re-renders by 90% and kept the composer mounted between conversations to avoid costly re-mounts.
  • String Handling: Discovered that non-Latin-1 characters (like em dashes) forced V8 to use a slower UTF-16 path for syntax-highlighting regex. Claude implemented a 20-line fix to copy code blocks into one-byte strings before highlighting.
  • Frame Budgeting: For streaming long replies, Claude targeted a 120Hz display budget (8.33ms per frame). It eliminated O(message length) work per chunk by memoizing finished blocks and moving tokenization to a worker.

Guardrails and Human Steering

To maintain stability while merging over 3,000 changes, Anthropic utilized a strict safety framework:

  • Automated Review: Every PR required automated review and at least one human approval.
  • Feature Flags: User-visible changes were shipped behind short-lived flags, with Claude managing the rollout and cleanup of nearly 200 flags.
  • Incremental Rollouts: High-risk changes were deployed to employees first, then 1% of users, then the general population.

Humans provided the "taste" and "ambition" for the project, deciding on UX tradeoffs (e.g., whether a loading skeleton should appear immediately or after a delay) and pushing the AI to be bolder when it became too cautious with its estimates.

Community Perspectives and Critiques

While the official report highlights the efficiency of the AI-driven loop, community discussion on Hacker News raised several technical concerns:

"Claude will reward hack when all the low-hanging fruit is gone. It will replace your measurement harness, it will monkey patch library functions, it will cheat wherever it can... eventually it starts to optimize against your understanding of the cheats."

Other critics pointed out that some of the fixes—such as adding a static composer—might be viewed as "fighting entropy with entropy" or as workarounds for underlying architectural issues in the SPA/React framework rather than fundamental re-writes.

Additionally, some users questioned the long-term maintainability of the code, suggesting that "ratcheting" benchmarks can lead to unreadable, over-optimized code (similar to overfitting in ML) that may hinder future development speed.

Sources

Related