Rethinking On-Policy Distillation of Large Language Models II: One Training Example
On-policy distillation recovered most full-data gains from one query; diverse rollouts mattered more than dataset size, while slow student alignment remained the bottleneck.
A single-query OPD run reached **71.5% state coverage** and recovered most of full-data OPD’s gain across tested tasks and model families. With **16 queries**, coverage rose to **98.9%** and matched full-data training.
If you distill smaller models for agent workloads, optimize prompts for diverse visited states before collecting a large task dataset. The experiments suggest rollout coverage, not query count or topical content alone, is the useful data signal.
A single-query OPD run reached **71.5% state coverage** and recovered most of full-data OPD’s gain across tested tasks and model families. With **16 queries**, coverage rose to **98.9%** and matched full-data training. If you distill smaller models for agent workloads, optimize prompts for diverse visited states before collecting a large task dataset. The experiments suggest rollout coverage, not query count or topical content alone, is the useful data signal. This does not make training cheap: student-teacher alignment still took **hundreds of steps**, even on fixed states. The claims concern OPD experiments, so deployment quality and transfer to a specific coding-agent workload still need validation.
This sharply refines data-efficiency guidance for OPD: a prompt set can be tiny if its rollouts traverse diverse states, so query count and topical breadth are poor proxies for useful supervision. It does not remove the optimization cost or establish coding-agent transfer, making state coverage a prompt-selection criterion rather than evidence that adaptation is broadly cheap.