Xiaomi Mimo 2.6 Post-Training Dashboard Analysis
Xiaomi has launched a live post-training dashboard for the Mimo 2.6 model series, offering an unprecedented level of transparency into the Reinforcement Learning (RL) process. The dashboard provides real-time telemetry on training costs, token throughput, and benchmark scores for two model variants: mimo-v2.6-pro and mimo-v2.6-flash.
Real-Time Training Metrics and Costs
Xiaomi is providing a granular look at the financial and computational investment required for post-training. As of the latest dashboard snapshot, the cumulative costs for the two runs are as follows:
- mimo-v2.6-pro: $740,289
- mimo-v2.6-flash: $321,466
Combined, the total run cost has exceeded $1.06 million. The training process involves massive token volumes, with the "flash" variant having processed 37.8B total tokens compared to 20.6B for the "pro" variant.
Performance Benchmarks: DeepSWE v1.1
The dashboard tracks performance on the DeepSWE v1.1 (mini-swe-agent, avg@3) benchmark, showing significant progress over previous versions.
- mimo-v2.6-pro: 62.24
- mimo-v2.6-flash: 60.77
For context, community discussions highlight that Mimo v2.5 Pro previously scored 19% on DeepSWE 1.1, indicating a substantial leap in coding capabilities for the 2.6 series. This puts the 2.6 models in a performance tier closer to other frontier coding agents like Fable (70%), Kimi K3 (69%), and Astra (74%).
Training Data Composition and RL Process
The RL training for Mimo 2.6 is heavily weighted toward software engineering and coding tasks. In step 10 of the pro run, the prompt distribution was:
- Code: 67.7% (1,062 prompts)
- Visual: 13.1% (206 prompts)
- General: 12.1% (190 prompts)
- Cyber: 4.1% (64 prompts)
- Chat: 3.0% (47 prompts)
Technical telemetry indicates a complex RL loop involving "rollouts," "judging," and "rewarding." The dashboard tracks metrics such as actor/entropy_loss, actor/pg_loss, and train_infer_diff/new_infer/kl to monitor model stability and divergence during the training steps.
Community Insights and User Experience
Users reporting experience with the Mimo series emphasize a high return on investment (ROI) due to low cost and high intelligence.
"The model is very powerful! Not perfect – I've run in hallucination loops once or twice... The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models."
Developers transitioning from v2.5 to the newer iterations note a marked improvement in multi-tasking and design capabilities. While v2.5-pro was described as a "conservative" coder that implemented minimal solutions, the newer versions are described as more "ambitious," making better guesses about gaps and next steps in a project.
Technical Observations and Controversies
The openness of the dashboard has sparked debate among technical observers regarding training methodologies:
- Contamination: Some users questioned whether running benchmarks during training constitutes data contamination.
- Distillation: There is speculation that the models are being distilled from other frontier models, with some users suggesting the use of Claude for generating training data.
- Transparency: The community has largely praised the move toward transparency, contrasting it with the more opaque training processes of US-based labs like OpenAI and Anthropic.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch