Surya is streaming RL training runs
Surya says their reinforcement learning training runs are being streamed at brusharena.art/live. The stream is inspired by LuoFuli.
Surya says their reinforcement learning training runs are being streamed at brusharena.art/live. The stream is inspired by LuoFuli.
OpenAI has published a page on how it thinks about securing frontier reinforcement-learning training runs, at a URL about safety cases for frontier AI training.
A paper from Amazon AGI proposes AutoGym, which writes an agent task, an executable environment, and a verifier together from a small domain seed or past model trajectories.
DAIR.AI describes a Salesforce AI Research method that isolates the multi-turn tool call whose action changes the later reward, then updates only that call. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points.
SmolDataEnvs is an open-source set of more than 5,000 verifiable reinforcement-learning environment tasks for hill-climbing small models in code and data science. The release includes environments, evaluations, and training.
Salesforce AI Research put that collection at 35.8% clean and two others at 10.1% and 3.3%, after finding reward errors in both directions. The authors say RIVER, which filters defective environments, increases RL gains by 106% on Terminal-Bench-Lite and 30% on Terminal-Bench v2.1 while using fewer than 30% of TMax's environments.