Critical-State RL trains only the tool call that changes the outcome
DAIR.AI describes a Salesforce AI Research method that isolates the multi-turn tool call whose action changes the later reward, then updates only that call. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points.