Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

CongoNames Corpus: A Large-Scale Labeled Dataset of Congolese Personal Names

Domaine:

natural language processing

Type de record:

dataset
Créateur:
TshAmaMer
Éditeur:
Zenodo
Hôte:avatar

Personal names carry cultural and linguistic identity, yet most African countries lack large-scale, structured name datasets suitable for natural language processing (NLP) research and computational social science. We present CONGONAMES, the first large-scale corpus of personal names from the Democratic Republic of the Congo (DRC), derived from publicly released national secondary-school examination palomàres (result lists) published annually by the DRC Ministry of Education. The corpus comprises 8,053,983 name records spanning 16 examination years (2008–2023) across 12 provinces and 304 sub-provincial regions, each enriched with a reported sex marker (M/F) and regional provenance metadata. We describe a fully deterministic, layered processing pipeline (bronze–silver–gold architecture) that converts raw PDF documents into structured CSV datasets without manual annotation or machine-learning-based inference. The dataset is validated against school-level census counts extracted from the same source PDFs, yielding extraction error rates below 2% for all years except 2023 (7.81%, flagged due to a layout change). Descriptive analyses document name length and token-count distributions, character-level n-gram profiles, provincial diversity indices, and inter-provincial name-inventory overlap, collectively establishing the dual linguistic origin—locally rooted Bantu components and Christian/French-origin components—that characterizes modern Congolese naming practice. The dataset, processing code, and documentation are released openly to support research in African NLP, onomastics, and computational social science.

Visit

doi.org

Languages

Ndasa

Tags

African LanguagesDemocratic Republic of CongoCongo NamesNatural Language ProcessingOnomastics

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

A Dataset of Personal Names in Gisamjanga DatoogaYpunu Personal Names Lexical DatasetThe Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African LanguagesHumBugDB: A Large-scale Acoustic Mosquito DatasetWAXAL: A Large-Scale Multilingual African Language Speech CorpusA subset of large-scale EEG dataset (India + Tanzania)

A Dataset of Personal Names in Gisamjanga Datooga

This dataset represents a collection of notes and recordings made in Haydom, Manyara, Tanzania by He

Ypunu Personal Names Lexical Dataset

The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages

We present the Thiomi Dataset, a large-scale multimodal corpus spanning ten African languages across

HumBugDB: A Large-scale Acoustic Mosquito Dataset

This paper presents the first large-scale multi-species dataset of acoustic recordings of mosquitoes

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech datase

A subset of large-scale EEG dataset (India + Tanzania)

Dataset ID: ds007358 Vianney2026 Canonical aliases: Vianney2025 At a glance: EEG · Resting State re