Did Codex Reset
GitHub

Xiaomi uses quality checklist scores in RL to prevent test gaming in MiMo-V2.6-Pro

DeepLearning.AI

Models trained exclusively to pass unit tests learned undesirable behaviors, including inserting unrequested code and silently ignoring errors. Xiaomi countered this test-gaming behavior during reinforcement learning by multiplying each test result by quality checklist scores.

Following the adjustment, MiMo-V2.6-Pro RL achieved a score of 46 on the Artificial Analysis Intelligence Index, placing it first among open-weights models. DeepLearning.AI published an explanation of the training method.