Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Small Languages, Big Models: A Study of Continual Training on Languages of Norway

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
SamMikVelØvr
Hôte:avatar
Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern Sámi. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokmål, Nynorsk, and Northern Sámi with 11.4 billion parameters: NorMistral-11B. Published at NoDaLiDa 2025

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and Language

Similaires

Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource LanguagesContinual-learning for Modelling Low-Resource Languages from Large Language ModelsImpact of Intermediate-Task Training on Low-Resource Languages in Multilingual Language ModelsSyntax-aware Offensive Content Detection in Low-resourced Code-mixed Languages with Continual Pre-trainingBig Math Translated - African LanguagesChallenges in Linguistic Research on African Languages: A Case Study of Nigerian Languages

Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages

Low-resource languages (LRLs) face significant challenges in natural language processing (NLP) due t

Continual-learning for Modelling Low-Resource Languages from Large Language Models

Modelling a language model for a multi-lingual scenario includes several potential challenges, among

Impact of Intermediate-Task Training on Low-Resource Languages in Multilingual Language Models

Accuracy of English-language Question Answering (QA) systems has improved significantly in recent ye

Syntax-aware Offensive Content Detection in Low-resourced Code-mixed Languages with Continual Pre-training

Social media is a widely used platform that includes a vast amount of user-generated content, allowi

Big Math Translated - African Languages

This is a set of 41k SynthLabsAI/Big-Math-RL-Verified questions translated into 9 African languages

Challenges in Linguistic Research on African Languages: A Case Study of Nigerian Languages

Abstract