Xiaomi Streams RL Training Process for MiMo-V2.6 Models
Xiaomi has begun a public livestream of its reinforcement learning training process for the MiMo-V2.6 family of models. The fully asynchronous pipeline processes 1,568 prompts per step with 16 rollouts each, handling approximately 2 billion tokens across mixed domains including code generation, cybersecurity, multimodal vision, and chat agent behavior. A public dashboard displays real-time metrics for both MiMo-v2.6-flash and MiMo-v2.6-pro, showing task distribution, context lengths, and step timings. Xiaomi plans to publish documentation and findings in open access in the coming weeks.