Did Codex Reset
GitHub

Harness-Aware Distillation trains small model agents to 63.4% success on ALFWorld

elvis

Researchers introduced Harness-Aware Distillation, a framework designed for small language model agents operating with an execution harness, according to an arXiv paper. The method queries the same teacher model with and without harness information, training the student model to prefer actions chosen when harness data is present. A filter discards candidate pairs where the preferred action contradicts harness records, operating without task rewards or success labels.

The researchers observed that simply adding harness data to on-policy distillation raised the student's harness usage on ALFWorld from 65.7% to 73.1%, but left task success flat at 43.1% to 43.5%. With Harness-Aware Distillation, the student model achieved a 63.4% success rate on unseen ALFWorld tasks, outperforming the best baseline at 47.0% and surpassing its 8B teacher. The trained student also escaped 59.7% of stalls, compared to baselines remaining near the untrained student baseline of 46.8%.