Did Codex Reset
GitHub

AI News

TGRSS feeds435 briefs

Only 35.8% of the cleanest public terminal-agent RL collection passed a Salesforce AI Research audit

Salesforce AI Research put that collection at 35.8% clean and two others at 10.1% and 3.3%, after finding reward errors in both directions. The authors say RIVER, which filters defective environments, increases RL gains by 106% on Terminal-Bench-Lite and 30% on Terminal-Bench v2.1 while using fewer than 30% of TMax's environments.

Artificial Analysis ranks Ming-Image-0.1-Design first among open-weight UI/UX models

On the Artificial Analysis Text to Image Leaderboard, the 6B Ming-Image-0.1-Design model places 17th of 81 in UI/UX design and first among open-weight models, compared with 45th of 160 overall. It is released under the MIT license with RGBA output and was evaluated at 2K resolution using 12 inference steps.