Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Corpus of English and Nigerian Pidgin Code-switching (CENCOS)

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AgbPla
Éditeur:
Zenodo
Hôte:avatar
This dataset was compiled from a fieldwork in Nigeria in 2019. It features naturally occurring spoken conversations from educated speakers of English and Nigerian Pidgin in Nigeria, with very few conversations involving uneducated speakers. Nigeria is a multi-lingual nation with over 500 languages that are not mutually intelligible. English and Nigerian Pidgin serve as lingua francas used to bridge linguistic gaps between speakers whose languages are mutually unintelligible. English is used in both formal and informal settings, but Nigerian Pidgin is used only in informal settings. Nigerian Pidgin was formerly regarded as the language of the uneducated in Nigeria. Over time, it has developed into a language spoken not only by the uneducated, but also by the educated in Nigeria. The compilation of this corpus is an effort to understand how the educated speakers with the knowledge of both languages are able to use them in interactions. This corpus contains both sound and text files, but the sound files are not included here for data protection reasons. The sound files are manually transcribed into texts, amounting to over 100, 000 word tokens. It contains no annotation other than the speakers. An excel sheet containing speakers’ basic information like gender, age, ethnic group and education status is included. With these social factors, this corpus is useful for any form of investigation on the use of English and Nigerian Pidgin in Nigeria.

Visit

doi.orgzenodo.org

Tasks

code switching

Tags

Nigerian PidginEducated speakersCode-switchingCorpus

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeOpen Accessinfo:eu-repo/semantics/openAccess

Similaires

Dataset on Teachers' Code Switching Between English and Nigerian Pidgin and Students' Classroom Participation in Rural Secondary SchoolsYoruba-English Code-Switching (YECS) Corpus | Mozilla Data CollectiveCode‐Switching: Amharic‐EnglishCorpora and Corpus Tools for Indigenous African Languages: The Case of an IsiZulu-English Spoken Word Code-Switching CorpusNigerian Pidgin English: morphology and syntaxKina-Research/English-Amharic-Code-Switching

Dataset on Teachers' Code Switching Between English and Nigerian Pidgin and Students' Classroom Participation in Rural Secondary Schools

This dataset contains a collected response sample of 600 Senior Secondary School students designed t

Yoruba-English Code-Switching (YECS) Corpus | Mozilla Data Collective

The Yoruba-English Code-Switching (YECS) Corpus is a comprehensive, ~120-hour dataset designed to capture the natural linguistic phenomenon of intra-sentential code-mixing. Curated by the LynguaTech Innovative Foundation (LyngualLabs), this dataset provides nearly

Code‐Switching: Amharic‐English

Corpora and Corpus Tools for Indigenous African Languages: The Case of an IsiZulu-English Spoken Word Code-Switching Corpus

Nigerian Pidgin English: morphology and syntax

Kina-Research/English-Amharic-Code-Switching

EthioSwitch-Bench Overview EthioSwitch-Bench is the first large-scale, human-validated benchmark