Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Datasets Creation and Empirical Evaluations of Cross-Lingual Learning on Extremely Low-Resource Languages : A Focus on Comorian Dialects

Domain:

natural language processing

Record type:

datasetpaper
Creator:
AssAbdou Mohamed Naira
Publisher:
Und
Host:avatar
In this era of extensive digitalization, there are a profusion of Intelligent Systems that attempt to understand how languages are structured for the aim of providing solutions in various tasks like Text Summarization, Sentiment Analysis, Speech Recognition, etc. But for multiple reasons going from lack of data to the nonexistence of initiatives, these applications are in an embryonic stage in certain languages and dialects, especially those spoken in the African continent, like Comorian dialects. Today, thanks to the improvement of Pre-trained Large Language Models, a spacious way is open to enable these kind of technologies on these languages. In this study, we are pioneering the representation of Comorian dialects in the field of Natural Language Processing (NLP) by constructing datasets (Lexicons, Speech Recognition and Raw Text datasets) that could be used on different tasks. We also measure the impact of using pre-trained models on languages closely related to Comorian dialects to enhance the state-of-the-art in NLP for these latter, compared to using pre-trained models on languages that may not necessarily be close to these dialects. We construct models covering the following use cases: Language Identification, Sentiment Analysis, Part-Of-Speech Tagging, and Speech Recognition. Ultimately, we hope that these solutions can catalyze the improvement of similar initiatives in Comorian dialects and in languages facing similar challenges.

Visit

doi.orgunderline.io

Tasks

automatic speech recognitionlanguage identificationsentiment analysisspeech processingtext classification

Languages

Comorian, Maore

Tags

Computational LinguisticsNatural Language Processing

Similar

Learning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel CorporaCross-Lingual Transfer Robustness to Lower-Resource Languages on Adversarial DatasetsImpact of Contrastive Learning on Cross-Lingual NER F1-Score in Low-Resource LanguagesCross-lingual transfer of multilingual models on low resource African LanguagesCross-lingual Transfer Effects on Euphemism Detection in Low-Resource LanguagesContrastive Learning Effects on XLM-R Zero-Shot Cross-Lingual Transfer in Low-Resource Languages

Learning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel Corpora

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small

Cross-Lingual Transfer Robustness to Lower-Resource Languages on Adversarial Datasets

Multilingual Language Models (MLLMs) exhibit robust cross-lingual transfer capabilities, or the abil

Impact of Contrastive Learning on Cross-Lingual NER F1-Score in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Cross-lingual transfer of multilingual models on low resource African Languages

Large multilingual models have significantly advanced natural language processing (NLP) research. Ho

Cross-lingual Transfer Effects on Euphemism Detection in Low-Resource Languages

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

Contrastive Learning Effects on XLM-R Zero-Shot Cross-Lingual Transfer in Low-Resource Languages

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia