Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

From N-grams to Pre-trained Multilingual Models For Language Identification

Domain:

natural language processing

Record type:

paperdataset
Creator:
Sindane, ThapeloMarivate, Vukosi
Host:avatar
In this paper, we investigate the use of N-gram models and Large Pre-trained Multilingual models for Language Identification (LID) across 11 South African languages. For N-gram models, this study shows that effective data size selection remains crucial for establishing effective frequency distributions of the target languages, that efficiently model each language, thus, improving language ranking. For pre-trained multilingual models, we conduct extensive experiments covering a diverse set of massively pre-trained multilingual (PLM) models -- mBERT, RemBERT, XLM-r, and Afri-centric multilingual models -- AfriBERTa, Afro-XLMr, AfroLM, and Serengeti. We further compare these models with available large-scale Language Identification tools: Compact Language Detector v3 (CLD V3), AfroLID, GlotLID, and OpenLID to highlight the importance of focused-based LID. From these, we show that Serengeti is a superior model across models: N-grams to Transformers on average. Moreover, we propose a lightweight BERT-based LID model (za_BERT_lid) trained with NHCLT + Vukzenzele corpus, which performs on par with our best-performing Afri-centric models. The paper has been accepted at The 4th International Conference on Natural Language Processing for Digital Humanities (NLP4DH 2024)

Visit

arxiv.org

Tasks

language identification

Tags

Computation and LanguageArtificial Intelligence

Similar

How Linguistically Fair Are Multilingual Pre-Trained Language Models?Adapting Pre-trained Language Models to African Languages via Multilingual Adaptive Fine-TuningGeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsPerformance Variation in Multilingual Pre-trained Language Models with Optimal Transport Distillation for AdversarialAfrikaans Literary Genre Recognition using Embeddings and Pre-Trained Multilingual Language ModelsIncreasing linguistic diversity in NLP : Fine-tuning Multilingual Pre-trained African Language Models

How Linguistically Fair Are Multilingual Pre-Trained Language Models?

Massively multilingual pre-trained language models, such as mBERT and XLM-RoBERTa, have received sig

Adapting Pre-trained Language Models to African Languages via Multilingual Adaptive Fine-Tuning

Multilingual pre-trained language models (PLMs) have demonstrated impressive performance on several downstream tasks for both high-resourced and low-resourced languages. However, there is still a large performance drop for languages unseen during pre-training, espe

GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models

Recent work has shown that Pre-trained Language Models (PLMs) store the relational knowledge learned

Performance Variation in Multilingual Pre-trained Language Models with Optimal Transport Distillation for Adversarial

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Afrikaans Literary Genre Recognition using Embeddings and Pre-Trained Multilingual Language Models

Increasing linguistic diversity in NLP : Fine-tuning Multilingual Pre-trained African Language Models

Increasing linguistic diversity in NLP : Fine-tuning Multilingual Pre-trained African Language Models

Poster presented at the Deep Learning Indaba 2023 by Fiskani Banda