Tencent Hy extends critical-batch-size theory to online LLM RL
Tencent Hy’s research on online LLM reinforcement learning revisits classical critical-batch-size theory and extends it to settings where the model generates its own training data, and where rollout generation and training scale differently.
Across GRPO and PPO, Tencent Hy finds that learning-rate retuning can preserve learning per response over a bounded range of batch sizes. On fixed hardware, scaling up the batch size improves PPO generation-stage throughput by up to 2.29×, while its best measured GRPO configuration reaches the same validation target in 29% less time.