RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer
RP-OPSD targets cross-lingual distillation at tokens that steer reasoning rather than surface wording, a useful training pattern for multilingual reasoning models.
**RP-OPSD** identifies reasoning pivots by comparing matched teacher views with and without an English reference solution. It was evaluated on mathematical reasoning benchmarks spanning **17 languages** and multiple difficulty levels.
If training multilingual reasoning models, consider weighting supervision toward reasoning-control decisions and state updates instead of treating every generated token equally. The method also uses reference anchoring around those pivots.
**RP-OPSD** identifies reasoning pivots by comparing matched teacher views with and without an English reference solution. It was evaluated on mathematical reasoning benchmarks spanning **17 languages** and multiple difficulty levels. If training multilingual reasoning models, consider weighting supervision toward reasoning-control decisions and state updates instead of treating every generated token equally. The method also uses reference anchoring around those pivots. The material reports gains over multilingual baselines and OPSD variants but provides no scores or model-level breakdowns. Evidence is limited to mathematical reasoning, so transfer to coding agents is still open.
This further narrows on-policy self-distillation from supervising all tokens to emphasizing reasoning-control pivots, with English-reference views anchoring transfer across 17 languages. Against the supplied distillation methods, it adds a multilingual criterion for locating valuable supervision rather than establishing a generally superior recipe. Missing scores and model breakdowns leave its advantage and transfer beyond mathematics unresolved.