Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Crosslingual On-Policy Self-Distillation for Multilingual Reasoning

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
LiuZhaHedSch
Hôte:avatar
Large language models (LLMs) have achieved remarkable progress in mathematical reasoning, but this ability is not equally accessible across languages. Especially low-resource languages exhibit much lower reasoning performance. To address this, we propose Crosslingual On-Policy Self-Distillation (COPSD), which transfers a model's own high-resource reasoning behavior to low-resource languages. COPSD uses the same model as student and teacher: the student sees only the low-resource problem, while the teacher receives privileged crosslingual context, including the problem translation and reference solution in English. Training minimizes full-distribution token-level divergence on the student's own rollouts, providing dense supervision while avoiding the sparsity and instability of outcome-only reinforcement learning (RL). Experiments on 17 low-resource African languages show that COPSD consistently improves low-resource mathematical reasoning across model sizes and substantially outperforms Group Relative Policy Optimization (GRPO). Further analyses show that COPSD improves answer-format adherence, strengthens test-time scaling, and generalizes to harder multilingual reasoning benchmarks, with especially large gains for lower-resource languages. We make our code and data available at: github.com. preprint

Visit

arxiv.org

Tags

Computation and Language

Similaires

Optimal Transport Distillation for Zero-Shot Cross-Lingual Reasoning in Multilingual LLMsCrosslingual Reasoning through Test-Time ScalingAlign to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math ReasoningSelf-Distillation of XLM-RoBERTa for Multilingual Sentiment Analysis on Imbalanced Data: A Case Study on Malagasy Mobile Application ReviewsImproving Multilingual Math Reasoning for African LanguagesLearning When to Translate for Multilingual Reasoning

Optimal Transport Distillation for Zero-Shot Cross-Lingual Reasoning in Multilingual LLMs

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Crosslingual Reasoning through Test-Time Scaling

Reasoning capabilities of large language models are primarily studied for English, even when pretrai

Align to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math Reasoning

Despite the impressive reasoning abilities demonstrated by large language models (LLMs), empirical e

Self-Distillation of XLM-RoBERTa for Multilingual Sentiment Analysis on Imbalanced Data: A Case Study on Malagasy Mobile Application Reviews

Improving Multilingual Math Reasoning for African Languages

Researchers working on low-resource languages face persistent challenges due to limited data availab

Learning When to Translate for Multilingual Reasoning

Reasoning language models (RLMs) achieve strong performance on complex reasoning tasks, but still ex