Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Collaborative construction of lexicographic and parallel datasets for African languages: first assessment

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
Tch
Éditeur:
arXiv
Hôte:avatar
Faced with a considerable lack of resources in African languages to carry out work in Natural Language Processing (NLP), Natural Language Understanding (NLU) and artificial intelligence, the research teams of NTeALan association has set itself the objective of building open-source platforms for the collaborative construction of lexicographic data in African languages. In this article, we present our first reports after 2 years of collaborative construction of lexicographic resources useful for African NLP tools. EACL 2021 - AfricaNLP workshop

Visit

doi.orgarxiv.org

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similaires

Constructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local LanguagesWebCrawl African : A Multilingual Parallel Corpora for African LanguagesOrature and morpholexical deconstruction as lexicographic archaeological sites: some implications for dictionaries of African languagesNatural Language Understanding Datasets for African LanguagesMultilingual Parallel Text Corpora for East African LanguagesAfri Code Datasets (A collection of datasets for code generation in African languages)

Constructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local Languages

In Indonesia, local languages play an integral role in the culture. However, the available Indonesia

WebCrawl African : A Multilingual Parallel Corpora for African Languages

WebCrawl African is a mixed domain multilingual parallel corpora for a pool of African languages com

Orature and morpholexical deconstruction as lexicographic archaeological sites: some implications for dictionaries of African languages

Natural Language Understanding Datasets for African Languages

Natural Language Understanding (NLU)

Natural Language Understanding (NLU) is a fundamental building block of goal-oriented dialogue systems like Alexa, Siri, Google Assistant, and Cortana (Fig 1). One of the major challenges of NLU is predicting the use

Multilingual Parallel Text Corpora for East African Languages

This is a partial multilingual parallel corpora of 5 East African languages. The dataset contains an

Afri Code Datasets (A collection of datasets for code generation in African languages)

Training and evaluating Large Language Models (LLMs) for code generation, building AI-powered coding