Weave Router 2.0: Open-Source Model Routing for Coding Agents
Weave Router 2.0 achieves performance parity with frontier models like GPT-6 Astra on key coding benchmarks while significantly reducing operational costs and latency. By intelligently switching between LLMs based on task complexity—using high-capability models for complex system design and lightweight models for simple updates—the router optimizes the trade-off between accuracy and efficiency.
Performance Benchmarks
Weave Router 2.0 was tested against GPT-6 Astra using Terminal Bench 4.0 and SWE Atlas. The results demonstrate that the ensemble approach matches the pass rates of a single frontier model while providing substantial efficiency gains:
| Benchmark | Pass Rate | Cost vs. Astra | Speed vs. Astra |
|---|---|---|---|
| Terminal Bench 4.0 | Equivalent | 52% | 2.2x faster |
| SWE Atlas | Equivalent | 54% | 2.5x faster |
Technical Architecture
Routing decisions in coding agent sessions are complex because a typical session with 100 turns and 10 available models creates a massive search space of potential paths. To manage this, Weave Router 2.0 employs three primary technical improvements:
Hidden Markov Model (HMM) and Classifier
To avoid the prohibitive cost of fully exploring the routing space via Reinforcement Learning (RL), the system uses a Hidden Markov Model to trace the session state. This HMM allows the router to understand not just the current state of the session, but the historical trajectory of how the session reached that state. A classifier then maps this session state to specific "buckets" of similar models, effectively pruning the search space by eliminating models that are unlikely to be effective for the given context.
Synthetic Data Bootstrapping
The router utilizes frontier LLMs to label a larger and more diverse set of coding agent sessions. This synthetic data is used to bootstrap both the HMM and the classifier, providing richer reward signals for the RL components of the routing logic.
Cache-Eviction Impact Calculation
To minimize costs, the router includes a subsystem that calculates the expected value of switching models. Because switching models often incurs a high one-time cost to refill the prompt cache for the new model, the router only triggers a switch when the predicted benefit of using a more capable model outweighs the cost of cache eviction.
Community Insights and Considerations
Following the announcement, the developer community raised several technical considerations regarding the implementation and evaluation of model routers:
- Performance Ceiling: Some users questioned whether a router trained on data labeled by frontier models can ever exceed the performance of those same models, suggesting that the primary gain is cost reduction rather than an increase in raw capability.
- Comparison to Existing Tooling: Users suggested benchmarking the router against other multi-model systems, such as Claude Code's "advisor mode" (Sonnet 5.5 with a Fable advisor) and Copilot's "HydraFusion" system, particularly for architecture planning and code review.
- State Tracking: There was discussion regarding the mathematical representation of the search space, with some noting that routing decisions are typically sequential (turn-by-turn) rather than a pre-determined path of 100 turns.
- Cache Efficiency: Concerns were raised about the potential for increased non-cache input tokens when switching models frequently, which could impact the overall cost-saving claims.
Weave Router is available as an open-source project on GitHub and as a hosted service.
Sources
Related
- Dispatch
- Project
- Dispatch
- Dispatch
- Dispatch