Sign InOpen Brain
arXivPaperNeeds Review

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

Small proxy models may be enough to choose an SFT-versus-RL annotation split: the paper finds broad near-optimal ranges that transfer to larger models.

arXiv · Sep 1, 2026
Open Source Open MarkdownOpen JSON
Source Summary

The study replaces a single best SFT–RL split with a range of allocations near peak performance. These regions remained wide at **2–10% tolerance**, widened with model scale, and transferred from small proxies to larger targets.

Practical Implication

If you post-train a model for agent workloads, search the allocation space on a small proxy first, then carry the near-optimal range into the expensive run. Budget using annotation cost, because unequal **SFT and RL data costs** shift that range.

Agent-Ready Context
The study replaces a single best SFT–RL split with a range of allocations near peak performance. These regions remained wide at **2–10% tolerance**, widened with model scale, and transferred from small proxies to larger targets.

If you post-train a model for agent workloads, search the allocation space on a small proxy first, then carry the near-optimal range into the expensive run. Budget using annotation cost, because unequal **SFT and RL data costs** shift that range.

The material reports consistency across tasks, model families, and **off-policy and on-policy RL**, but gives no concrete savings or universal allocation ratio. Transfer still needs validation for a builder’s own data and evaluation target.
Connected Context · Feed7 Judgment

This changes post-training allocation from a search for one universal SFT–RL ratio into finding a cost-aware region that remains near peak performance. It offers a practical proxy-to-target search strategy and suggests larger models may tolerate broader allocations, but does not remove workload validation or establish a universal split, savings level, or stopping rule.

Adaption Labs: Gradient-Free Continual Learning — Sara Hooker, AdaptionAuto Scientist proposes broader automation of training choices; this work supplies a bounded procedure for one such choice—search the SFT–RL allocation region on a small proxy before scaling.Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAIData-quality work argues that training value depends on curation and task fit, reinforcing why allocation should use annotation cost and evaluated performance rather than raw SFT and RL example counts.When Does Bigger Help? A Controlled Study of LLM Scale for Ontology LearningThe controlled scale study cautions that size alone does not predict task performance; similarly, proxy-transferred allocation ranges still require validation against the target model, data, and evaluation.
Context Map
modeldata#model-selection
Uncertainty
The material reports consistency across tasks, model families, and **off-policy and on-policy RL**, but gives no concrete savings or universal allocation ratio. Transfer still needs validation for a builder’s own data and evaluation target.