Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language

Domaine:

natural language processing

Type de record:

paperdatasetmodelsoftware
Créateur:
KrsTasSazGjoreski, Hristijan
Hôte:avatar
The increase in technological adoption worldwide comes with demands for novel tools to be used by the general population. Large Language Models (LLMs) provide a great opportunity in this respect, but their capabilities remain limited for low-resource languages, restricting applications in countries where such languages are spoken. We create several resources to facilitate the adoption of LLMs and to support research advancements for Macedonian. We collect the largest Macedonian corpus to date, consisting of 40GB of textual data and totaling 3.5B words. To support conversational applications, we collect a 106k-instance instruction dataset, carefully built to be culturally grounded. For evaluation, we construct a Macedonian evaluation suite covering seven benchmarks. Finally, we train domestic-yak, a state-of-the-art 8B-parameter model, on our curated datasets and evaluate it against eight baseline models using the newly constructed benchmark suite. Our model outperforms all existing models in the 8B parameter range across all benchmarks, and achieves performance comparable to models up to 10x larger. Furthermore, a qualitative analysis with native speakers reveals that our model is preferred over larger counterparts, receiving higher ratings for grammatical correctness and cultural appropriateness. All datasets, code, and model weights are openly released, setting a foundation for advancing LLMs in similarly underrepresented languages. These resources are publicly available at github.com for source code, and at huggingface.co for pretrained model weights and data. Camera-ready version accepted at SlavNLP-2025@ACL

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and Language

Similaires

Foundation Models for Low-Resource Language Education (Vision Paper)A Very Low Resource Language Speech Corpus for Computational Language Documentation ExperimentsMaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili LanguageLow-Resource Corpus Indonesian Local LanguagePashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource LanguageInkubaLM: A small language model for low-resource African languages

Foundation Models for Low-Resource Language Education (Vision Paper)

Recent studies show that large language models (LLMs) are powerful tools for working with natural la

A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments

Most speech and language technologies are trained with massive amounts of speech and text informatio

MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language

Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due

Low-Resource Corpus Indonesian Local Language

This study departs from the hypothesis that combining Neural Machine Translation (NMT) with the stem

Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language

We present the Pashto Common Voice corpus -- the first large-scale, openly licensed speech resource

InkubaLM: A small language model for low-resource African languages

High-resource language models often fall short in the African context, where there is a critical nee