Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Embedding model for the Malagasy informal language

Domaine:

natural language processing

Type de record:

model
Créateur:
FraAimNda
Éditeur:
Uni
Hôte:
Processing informal Malagasy language presents major challenges due to linguistic variations, abbreviations, and frequent code-switching in digital communication. This study proposes a text embedding model based on DistilBERT and XML-RoBERTa, specifically adapted to informal Malagasy. Through fine-tuning on custom corpora, we observe a gradual improvement in performance, with a significant reduction in loss function and lower perplexity, indicating a better understanding of linguistic structures. The evaluation shows that the generated embeddings effectively capture semantic similarities, even across varied formulations. DistilBERT outperforms XML-RoBERTa, demonstrating better generalization. These results highlight the importance of adapting language processing models to low-resource languages and open up new perspectives for applications in the automatic understanding of informal language.

Visit

doi.org

Tasks

embeddings

Languages

MalagasyMalagasy, Merina