Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

$M^3$ Scaling Law: Optimizing Multi-Epoch, Multi-Lingual, and Multi-Stage Training for Low-Resource Language Models

Domaine:

natural language processing

Type de record:

paper
Créateur:
AkiMiyOya
Hôte:avatar
In this paper, we study a fundamental design problem in pretraining Large Language Models (LLMs) for low-resource language regimes. Existing works adopt multi-epoch, multi-lingual, and multi-stage training to utilize the limited target-language corpus efficiently, but no prior scaling law can compare recipes spanning these approaches under the same compute budget $C$ and target-language corpus size $D_T$, leaving the optimal training setup unclear. To address this gap, we propose the $M^3$ Scaling Law, a unified predictive model parameterized by the model scale, the number of target-corpus epochs $k$, the average target-language ratio $r$, and the final-stage target-language ratio $r_f$, which places monolingual single-stage, multi-lingual single-stage, and multi-lingual multi-stage recipes on a single target-language loss surface. Across three language pairs, it extrapolates to unseen hyperparameter regions more accurately than existing scaling laws. Using $M^3$ as a surrogate objective, we derive two practical guidelines for low-resource LLM pretraining: (i) as $D_T$ decreases, the optimal recipe shifts directly from monolingual single-stage to multi-lingual two-stage training at a compute-budget-dependent threshold, with multi-lingual single-stage never optimal in our experimental grid; and (ii) the optimal number of epochs collapses onto a single curve in the scarcity variable $D_T/D^*(C)$, where $D^*(C) \propto C^{α/(α+β)}$ is the monolingual compute-optimal corpus size. 35 pages, 14 figures, 17 tables

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and Language

Similaires

Scaling Source Language Diversity in Multi-Source Cross-Lingual NER for Low-Resource WikiAnn PerformanceSelf-training in Multi-source Cross-lingual NER for Low-resource LanguagesMulti-lingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer in Low-Resource LanguagesMulti-source Intermediate-task Training for Low-resource XTREME Language GeneralizationMulti-Source Teacher-Student Models for Cross-Lingual NER in Low-Resource LanguagesMulti-lingual Functional Evaluation for Large Language Models

Scaling Source Language Diversity in Multi-Source Cross-Lingual NER for Low-Resource WikiAnn Performance

To better tackle the named entity recognition (NER) problem on languages with little/no labeled data

Self-training in Multi-source Cross-lingual NER for Low-resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Multi-lingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Multi-source Intermediate-task Training for Low-resource XTREME Language Generalization

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Multi-Source Teacher-Student Models for Cross-Lingual NER in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Multi-lingual Functional Evaluation for Large Language Models

Multi-lingual competence in large language models is often evaluated via static data benchmarks such