Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Natural Language Understanding Datasets for African Languages

Domaine:

natural language processing

Type de record:

posterdatasetproject

Natural Language Understanding (NLU)

Natural Language Understanding (NLU) is a fundamental building block of goal-oriented dialogue systems like Alexa, Siri, Google Assistant, and Cortana (Fig 1). One of the major challenges of NLU is predicting the user's intent and the slots (arguments) of that intent from their query as seen on the figure below. Several NLU resources exist for high resource languages like English, however, African languages, which constitute over 2000 spoken languages globally, are underrepresented in NLU resources and systems. To address this, we extend the ATIS dataset to 3 African languages: Swahili, Luganda, and Kinyarwanda as the first steps towards improving the representation of African Languages in goal-oriented dialogue systems.

PROBLEM STATEMENT

The limited availability of Natural Language Understanding (NLU) resources for Low Resource Languages (LRLs), particularly African languages which hinders the development of goal-oriented conversational AI systems for these languages.

STUDY OBJECTIVES

  • Creating re-usable translation pipelines for African languages
  • Creating and open source NLU datasets for African Languages to enable further development in the domain
  • Train deep learning models under semi-supervised and unsupervised experiments
  • Publish a research paper

DATASET

We used the ATIS dataset, comprising of 4978 training and 893 test samples. We translated the utterance into Swahili and Kinyarwanda using Google Translation API, and hired native translators to correct them. Only the test set was annotated for intents and slots in Kinyarwanda and Swahili using a custom annotation tool(scan annotation tool which the GitHub link is available in the uploaded datasets).

Visit

drive.google.comaccess datasetproject

Languages

KinyarwandaSwahili

Tags

posterdeep learning indaba 2023deep learning indabadlinatural language understanding for african languages

Similaires

AimeM250/Natural-Language-Understanding-System-for-African-LanguagesNatural language processing for African languages NLP for African languagesNatural language processing for African languagesCheetah: Natural Language Generation for 517 African LanguagesNatural Language Processing for African Languages in Kenya: Challenges and OpportunitiesNatural Language Processing for African Languages in Botswana: Challenges and Opportunities

AimeM250/Natural-Language-Understanding-System-for-African-Languages

The following repository contains a Natural Language Understanding for African languages (NLU) proje

Natural language processing for African languages NLP for African languages

Natural language processing for African languages

Recent advances in pre-training of word embeddings and language models leverage large amounts of unlabelled texts and self-supervised learning to learn distributed representations that have significantly improved the performance of deep learning models on a large v

Cheetah: Natural Language Generation for 517 African Languages

Low-resource African languages pose unique challenges for natural language processing (NLP) tasks, i

Natural Language Processing for African Languages in Kenya: Challenges and Opportunities

Natural Language Processing (NLP) has seen significant advancements in English and other ma

Natural Language Processing for African Languages in Botswana: Challenges and Opportunities

{ "background": "Natural Language Processing (NLP) is a critical field within Computer Scie