Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Extending new language in NLLB-200: language informal Malagasy

Domain:

natural language processing

Record type:

model
Creator:
FraAimNda
Publisher:
Uni
Host:
This study focuses on integrating informal Malagasy into the NLLB-200 model for machine translation. The model underwent supervised pretraining, which quickly led to improved performance, marked by a significant reduction in both loss and perplexity. This step allowed the model to effectively adapt to the unique linguistic structures of Malagasy. The evaluation of key translation metrics such as BLEU, ROUGE, and BertScore showed that the model produces high-quality translations, combining fluency with semantic coherence. Although the BLEU score was moderate, the ROUGE and BertScore results revealed a remarkable level of lexical and semantic fidelity. This work highlights the importance of developing translation systems that can handle low-resource languages, which are often overlooked by traditional technologies. The study also demonstrates the model’s ability to grasp the nuances of informal Malagasy, resulting in significant improvements over existing translation tools. In conclusion, this approach emphasizes the need to include informal languages in translation systems, paving the way for more inclusive and linguistically tailored applications.

Visit

doi.org

Tasks

machine translation

Languages

MalagasyMalagasy, Merina

Similar

Embedding model for the Malagasy informal languagemorlayecis0003/nllb-200-nllb-francais-pulaarAfri-XNLI+: Extending the XNLI Dataset with New African Language TranslationsSakuzas/nllb-200-wolaytta_to_english_kunllb-200-10M-samplemwkhettab/nllb-200-en-darija

Embedding model for the Malagasy informal language

Processing informal Malagasy language presents major challenges due to linguistic variations, abbrev

morlayecis0003/nllb-200-nllb-francais-pulaar

Afri-XNLI+: Extending the XNLI Dataset with New African Language Translations

Afri-XNLI+ is a multilingual dataset consisting of translated premise-hypothesis sentence pairs from

Sakuzas/nllb-200-wolaytta_to_english_ku

nllb-200-10M-sample

This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored w

mwkhettab/nllb-200-en-darija