Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

WolBanking77: Wolof Banking Speech Intent Classification Dataset

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
KanPreBa,Ndi
Éditeur:
UniUniScaMod
Éditeur:
CCSD
Hôte:avatar
International audience Intent classification models have made a significant progress in recent years. However, previous studies primarily focus on high-resource language datasets, which results in a gap for low-resource languages and for regions with high rates of illiteracy, where languages are more spoken than read or written. This is the case in Senegal, for example, where Wolof is spoken by around 90% of the population, while the national illiteracy rate remains at of 42%. Wolof is actually spoken by more than 10 million people in West African region. To address these limitations, we introduce the Wolof Banking Speech Intent Classification Dataset (WolBanking77), for academic research in intent classification. WolBanking77 currently contains 9,791 text sentences in the banking domain and more than 4 hours of spoken sentences. Experiments on various baselines are conducted in this work, including text and voice state-of-the-art models. The results are very promising on this current dataset. In addition, this paper presents an in-depth examination of the dataset’s contents. We report baseline F1-scores and word error rates metrics respectively on NLP and ASR models trained on WolBanking77 dataset and also comparisons between models. Dataset and code available at: wolbanking77.

Visit

hal.science

Tasks

automatic speech recognitionspeech processingtext classification

Languages

Wolof

Tags

low-resource languagesmultilingual speech recognitionafrican languageswolofnatural language processingautomatic speech recognitionintent classification[INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similaires

KanAgriIntent-7200: A Multi-Script Kannada Agricultural Intent Classification DatasetMultilingual Speech Intent Recognition Dataset for Home Automation: Luganda and RunyankoreHappymoreMasoka/shona-intent-classification-modelWolBanking77Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal Kallaama Wolof speech dataset Kallaama Pulaar speech dataset Kallaama Sereer speech datasetabdoukarim/wolbanking77

KanAgriIntent-7200: A Multi-Script Kannada Agricultural Intent Classification Dataset

KanAgriIntent-7200 is the first intent classification dataset for Kannada agricultural dialogue, com

Multilingual Speech Intent Recognition Dataset for Home Automation: Luganda and Runyankore

This dataset is a curated collection designed for multilingual speech intent recognition, specifical

HappymoreMasoka/shona-intent-classification-model

WolBanking77

An Intent Classification Dataset for Wolof language in Banking domain

Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal Kallaama Wolof speech dataset Kallaama Pulaar speech dataset Kallaama Sereer speech dataset

This data is transcribed speech data, in Wolof, Pulaar and Sereer. The recordings are about agricul

abdoukarim/wolbanking77

Wolof Banking Speech Intent Classification Dataset training and evaluation code. # Wolbanking77: Wo