Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
Small proxy models may be enough to choose an SFT-versus-RL annotation split: the paper finds broad near-optimal ranges that transfer to larger models.
The study replaces a single best SFT–RL split with a range of allocations near peak performance. These regions remained wide at **2–10% tolerance**, widened with model scale, and transferred from small proxies to larger targets.
If you post-train a model for agent workloads, search the allocation space on a small proxy first, then carry the near-optimal range into the expensive run. Budget using annotation cost, because unequal **SFT and RL data costs** shift that range.
The study replaces a single best SFT–RL split with a range of allocations near peak performance. These regions remained wide at **2–10% tolerance**, widened with model scale, and transferred from small proxies to larger targets. If you post-train a model for agent workloads, search the allocation space on a small proxy first, then carry the near-optimal range into the expensive run. Budget using annotation cost, because unequal **SFT and RL data costs** shift that range. The material reports consistency across tasks, model families, and **off-policy and on-policy RL**, but gives no concrete savings or universal allocation ratio. Transfer still needs validation for a builder’s own data and evaluation target.
This changes post-training allocation from a search for one universal SFT–RL ratio into finding a cost-aware region that remains near peak performance. It offers a practical proxy-to-target search strategy and suggests larger models may tolerate broader allocations, but does not remove workload validation or establish a universal split, savings level, or stopping rule.