Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Embedding model for the Malagasy informal language

Domain:

natural language processing

Record type:

model
Creator:
FraAimNda
Publisher:
Uni
Host:
Processing informal Malagasy language presents major challenges due to linguistic variations, abbreviations, and frequent code-switching in digital communication. This study proposes a text embedding model based on DistilBERT and XML-RoBERTa, specifically adapted to informal Malagasy. Through fine-tuning on custom corpora, we observe a gradual improvement in performance, with a significant reduction in loss function and lower perplexity, indicating a better understanding of linguistic structures. The evaluation shows that the generated embeddings effectively capture semantic similarities, even across varied formulations. DistilBERT outperforms XML-RoBERTa, demonstrating better generalization. These results highlight the importance of adapting language processing models to low-resource languages and open up new perspectives for applications in the automatic understanding of informal language.

Visit

doi.org

Tasks

embeddings

Languages

MalagasyMalagasy, Merina