Did Codex Reset
GitHub

#salesforce

View all AI news
TGRSS feeds2 briefs

Only 35.8% of the cleanest public terminal-agent RL collection passed a Salesforce AI Research audit

Salesforce AI Research put that collection at 35.8% clean and two others at 10.1% and 3.3%, after finding reward errors in both directions. The authors say RIVER, which filters defective environments, increases RL gains by 106% on Terminal-Bench-Lite and 30% on Terminal-Bench v2.1 while using fewer than 30% of TMax's environments.