Logo Lanfrica

Replication Data for Igbo Natural Language Processing Tasks II Igbo Synchronised Corpus for Natural Language Processing Tasks

Domaine:

natural language processing

Type de record:

dataset
Créateur:
NweAkiOnwEji
Éditeur:
LACNweNwe
Éditeur:
Har
Hôte:avatar
The Igbo synchronised corpus (IgboSynCorp) is an annotated corpus of spoken Igbo created by a team of linguists and NLP experts at the University of Ibadan and Afe Babalola University, Nigeria. The project was designed to create an open access labelled and unlabelled dataset for Natural Language Processing tasks in the Igbo language. The dataset was created to enable robust and more equitable application of machine learning tools of high social value in Igbo. The dataset is consists of ELAN text and wav files of Igbo speech. There are two categories of ELAN files: Gold files (90 mins) and Non Gold files (188 mins). The Gold files (19,722 words or 2761 sentences were transcribed phonetically and orthographically, translated to English, glossed and PoS tagged based on the universal dependency PoS tags . The None Gold files were only transcribed orthographically and translated to English. There are 110 recordings of spoken Igbo (.wav Files) amounting to 38.8075 hours or 2,328.45 minutes. There are 110 wav files of Igbo Oral narratives. The metadata is compiled in excel sheets. The Igbosyncorp Metadata I contains the demographic information about the language consultants. While Igbosyncorp metadata II outlines domains of speech represented in the individual wav file (oral narrative). There are two lexicon files with about 2300 words altogether which originated from the glossing and part of speech tagging, The project was funded by Lacuna Fund lacunafund.org of the Meridian Institute, 105 Village Place, Dillion, Colorado 80435, United States of America. Replication Data for Igbo Natural Language Processing Tasks I The wav files are labelled based on the five Southeast states of Nigeria where the oral narratives were collected. The states are Abia, Anambra Ebonyi, Enugu and Imo. Out of 105 wav files 13 have a corresponding eaf (ELAN) files. They are Abia_0002, 0004, 0005, 0010, Anambra_0002, 0010, 0011, Ebony_0011, 0018, Enugu_0014, 0025; Imo_0005, 0011. All the files were anonymised using audacity. We have consent from all the consultants to launch the dataset open access. See Igbosyncorp metadata I and II for the demographic information and speech domains.