Did Codex Reset
GitHub

Critical-State RL trains only the tool call that changes the outcome

DAIR.AI

Critical-State RL targets multi-turn tool use by training the single call where the action changes the outcome, rather than spreading reward across the whole trajectory. When the reward depends on later turns, much of its variation comes from what happens downstream.

The method uses nested sampling to separate reward variation caused by the current action from that downstream noise, then trains only the selected call with contextual-bandit updates. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points, while training the other candidate turn leaves accuracy flat or lower.

DAIR.AI points to the Critical-State RL paper.