Story thread · 2 reports / 2 sources
Xiaomi open-sources MiMo-V2.6, publishes the bill for scaling reinforcement learning
digitimes.com · 7h · first report

Chinese labs have spent the year arguing that reinforcement learning, not pre-training, holds the next gains. Xiaomi has now put a price on that argument. The company released and open-sourced its MiMo-V2.6 series on September 22, after livestreaming the reinforcement learning post-training run that produced it — publishing step counts, token consumption, and a running cost meter on a public page as the models trained.
The coverage
The conversation · 0
Sign in to join the conversation.
No comments yet — start the thread.