Tencent Hy extends critical-batch-size theory to online LLM RL
Tencent Hy says learning-rate retuning can preserve learning per response across GRPO and PPO over a bounded range of batch sizes. On fixed hardware, larger batches improve PPO generation-stage throughput by up to 2.29×, and its best measured GRPO configuration reaches the same validation target in 29% less time.