Olmo-core 3 Open Mixture-of-Experts Training Infrastructure
TL;DR
Olmo-core 3 is a newly released open‑source training stack that enables efficient mixture‑of‑experts (MoE) models at the trillion‑parameter scale, delivering up to 2.7× higher token throughput than the prior implementation and supporting advanced optimizations such as MXFP8 precision.
What Olmo‑core 3 Is and Why It Matters
Olmo‑core 3 is a redesign of the Allen Institute’s MoE training framework, built to keep computational efficiency as expert pools grow from dozens to hundreds while the number of active parameters per token stays roughly constant. By closing the communication and routing overhead that traditionally erodes MoE advantages, the stack makes trillion‑parameter sparse models affordable for academic labs and smaller research teams.
Core Architectural Changes
Switch from FSDP to DDP
- Earlier Olmo‑core used Fully Sharded Data Parallelism (FSDP) which repeatedly gathered and reshared model weights for each mini‑batch.
- Olmo‑core 3 adopts Distributed Data Parallelism (DDP), keeping expert weights resident on GPUs and routing only the necessary token data. This eliminates the costly weight‑gather step.
Integrated MoE Stack vs. Megatron‑Core
- NVIDIA’s Megatron‑Core remains a reference implementation for large MoEs.
- Olmo‑core 3 provides an end‑to‑end, tightly integrated MoE stack that outperforms the previous FSDP‑based version. In a benchmark on eight NVIDIA B300 GPUs, a 47 B‑parameter MoE achieved 52,000 tokens / s / GPU, compared with 19,400 tokens / s / GPU previously – a 2.7× throughput increase.
Scaling Techniques Employed
Parallelism Strategies
| Technique | Purpose |
|---|---|
| Expert parallelism | Distributes the expert pool across GPUs so each GPU stores only a subset of experts. |
| Pipeline parallelism | Splits model layers across GPU groups, reducing per‑GPU memory footprint. |
| Distributed optimizer | Spreads optimizer state across GPUs, avoiding a full copy on every device. |
These three mechanisms together allow MoE models to scale without requiring every GPU to hold the entire model or optimizer state.
Routing and Computation Optimizations
- Rowwise expert parallelism – Directly writes routed tokens into expert input buffers, cutting down on data reshaping.
- GPU‑resident routing – Keeps routing metadata on the GPU, so the CPU does not become a bottleneck.
- Grouped GEMM – Batches many small matrix multiplications from different experts into a single, larger GEMM call, improving GPU utilization.
Precision Optimization with MXFP8
- MXFP8 is a lower‑precision numeric format that reduces both compute and inter‑GPU traffic.
- In a controlled benchmark on four NVIDIA B300 GPUs, enabling MXFP8 raised end‑to‑end throughput by ~21 % over the BF16 baseline while lowering peak active memory from 103 GiB to 95 GiB.
- Most gains stemmed from faster feed‑forward computation and reduced expert‑to‑expert data movement.
Demonstrated Scale
| Configuration | Total parameters | Active parameters per token | GPUs | Peak throughput |
|---|---|---|---|---|
| 1.2 T‑parameter MoE | 1.2 T | 58.36 B | 512 | 858 TFLOP/s / GPU |
| 2.38 T‑parameter (DeepEP v2) | 2.38 T | — | — | — |
The 1.2 T‑parameter run used random routing to stress‑test system performance; it does not reflect a trained model’s quality.
Empirical Findings Beyond Raw Speed
- Token gerrymandering – A routing‑balance score can improve while actual workload balance degrades, highlighting a metric‑design pitfall.
- Expert learning‑rate scaling – Lowering learning rates for sparsely‑used experts did not yield better results in the tested model family.
- Input‑value dependent timing – GPU kernels exhibited variable runtimes for identical matrix shapes when input values differed, underscoring the need for value‑matched benchmarks.
- Communication‑computation overlap – Overlapping these streams sometimes slowed overall execution, showing that more overlap is not universally beneficial.
These observations are detailed in the accompanying technical report and guide future MoE training research.
Open‑Source Availability and Community Impact
- Code – Fully open on GitHub: https://github.com/allenai/olmo-core
- Technical report – Provides in‑depth system design, ablations, and experimental methodology: https://allenai.org/papers/olmocore3
- Interactive demo – Visual walkthrough of data, expert, and pipeline parallelism: https://narrative.allen.ai/scaling-up-training
By releasing the stack under an open license, Allen Institute enables researchers to train their own trillion‑parameter MoEs, adapt the system to alternative hardware, and experiment with novel routing or precision schemes.
Outlook
Olmo‑core 3 forms the backbone for the next generation of Olmo models, which will combine the new MoE architecture with larger datasets and longer context windows. The stack’s modular design ensures it can evolve alongside hardware advances, keeping the research community at the forefront of scalable sparse‑model training.
Key Takeaways
- Olmo‑core 3 scales MoE training to the trillion‑parameter regime while keeping per‑token active parameters constant.
- Through DDP, expert/pipeline parallelism, and routing optimizations, it achieves up to 2.7× higher token throughput than the prior FSDP‑based stack.
- MXFP8 precision adds a further ~21 % speed boost and reduces memory usage.
- The framework is fully open‑source, providing the community with a production‑grade tool for large‑scale sparse model research.