Did Codex Reset
GitHub

SkillGym fine-tunes Qwen3.5-35B-A3B on skill trajectories

DAIR.AI

A paper proposes SkillGym, which turns human-written skills into 2,756 training environments with code-based checkers and fine-tunes on 8,364 successful trajectories. The paper is SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving.

After fine-tuning, Qwen3.5-35B-A3B running in Claude Code gains 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%. The paper reports that result as above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model beats the base model that has the skills in context.