Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

GeezSwitch: Language Identification in Typologically Related Low-resourced East African Languages

Domain:

natural language processing

Record type:

paper
Language identification is one of the fundamental tasks in natural language processing that is a prerequisite to data processing and numerous applications. Low-resourced languages with similar typologies are generally confused with each other in real-world applications such as machine translation, affecting the user’s experience. In this work, we present a languageidentification dataset for five typologically and phylogenetically related low-resourced East African languages that use the Ge’ez script as a writing system; namely Amharic, Blin, Ge’ez, Tigre, and Tigrinya. The dataset is built automatically from selected data sources, but we also performed a manual evaluation to assess its quality. Our approach to constructing the dataset is cost-effective and applicable to other low-resource languages. We integrated the dataset into an existing language-identification tool and also fine-tuned several Transformer based language models, achieving very strong results in all cases. While the task of language identification is easy for the informed person, such datasets can make a difference in real-world deployments and also serve as part of a benchmark for language understanding in the target languages

Visit

www.lrec-conf.org

Connected records

project

Tasks

language identification

Languages

AmharicBilenGeezTigréTigrigna

Similar

TEXT-BASED LANGUAGE IDENTIFICATION FOR TYPOLOGICALLY RELATED ETHIOPIAN LANGUAGES: A DEEP LEARNING APPROACHShort Text Language Identification for Under Resourced LanguagesSurface Realization Architecture for Low-resourced African LanguagesAdvancing sentiment analysis for low-resourced african languages using pre-trained language modelsGlotLID: Language Identification for Low-Resource LanguagesBuilding Text and Speech Datasets for Low Resourced Languages: A Case of Languages in East Africa

TEXT-BASED LANGUAGE IDENTIFICATION FOR TYPOLOGICALLY RELATED ETHIOPIAN LANGUAGES: A DEEP LEARNING APPROACH

TEXT-BASED LANGUAGE IDENTIFICATION FOR TYPOLOGICALLY RELATED ETHIOPIAN LANGUAGES: A DEEP LEARNING AP

Short Text Language Identification for Under Resourced Languages

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text languag

Surface Realization Architecture for Low-resourced African Languages

There has been growing interest in building surface realization systems to support the automatic gen

Advancing sentiment analysis for low-resourced african languages using pre-trained language models

While sentiment analysis systems excel in high-resource languages, most African languages facing lim

GlotLID: Language Identification for Low-Resource Languages

International audience Several recent papers have published good solutions for langua

Building Text and Speech Datasets for Low Resourced Languages: A Case of Languages in East Africa

Africa has over 2000 languages; however, those languages are not well represented in the existing Natural Language Processing ecosystem. African languages lack essential digital resources to be engaged effectively in the advancing language technologies. This growin