Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Development of an annotated Yoruba text corpus for automatic event extraction

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AdeOlu
Éditeur:
Res
Hôte:
Abstract The study documents the development of an annotated Yoruba text corpus which serves as a fundamental reference for training and evaluating automatic event extraction models specifically designed for the Yoruba language. While notable corpora like ACE and ERE have been established for high-resource and extensively annotated languages, African languages such as Yoruba have faced limitations in information extraction due to the lack of annotated datasets for training and model evaluation. In this study, the researchers took the initiative to preprocess raw Yoruba text obtained from selected Yoruba folktale books, ensuring the correct placement of tone marks, and proceeded to annotate the data with events and their corresponding arguments, including temporal and spatial information, utilizing the widely recognized BIO annotation format. This meticulously developed corpus now fills a critical void and can serve as a solid foundation for any future research endeavours involving the Yoruba language in the domain of information extraction.

Visit

doi.org

Tasks

information extraction

Languages

Yoruba

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Development of an annotated Yoruba text corpus for automatic event extraction تطوير مجموعة نصوص اليوروبا المشروحة لاستخراج الحدث التلقائي Élaboration d'un corpus textuel yoruba annoté pour l'extraction automatique d'événements Desarrollo de un corpus de texto yoruba anotado para la extracción automática de eventosDataset for Event Extraction in Yoruba LanguageAUTOMATIC RELATION EXTRACTION BETWEEN ENTITIES FOR AMHARIC TEXTAlgBERT: Automatic Construction of Annotated Corpus for Sentiment Analysis in Algerian DialectPOS annotated corpus with 5 different text types for isiZuluLULCC-KnowText - annotated text segments for knowledge extraction on Land Use and Land Cover change

Development of an annotated Yoruba text corpus for automatic event extraction تطوير مجموعة نصوص اليوروبا المشروحة لاستخراج الحدث التلقائي Élaboration d'un corpus textuel yoruba annoté pour l'extraction automatique d'événements Desarrollo de un corpus de texto yoruba anotado para la extracción automática de eventos

Abstract The study documents the development of an annotated Yoruba text corpus which serves as a fu

Dataset for Event Extraction in Yoruba Language

Event Extraction Dataset

AUTOMATIC RELATION EXTRACTION BETWEEN ENTITIES FOR AMHARIC TEXT

This research work primarily focused on the automatic relation extraction between entities for Amhar

AlgBERT: Automatic Construction of Annotated Corpus for Sentiment Analysis in Algerian Dialect

Nowadays, sentiment analysis is one of the most crucial research fields of Natural Language Processi

POS annotated corpus with 5 different text types for isiZulu

This is a POS annotated corpus with 5 different text types for isiZulu. The text types included a

LULCC-KnowText - annotated text segments for knowledge extraction on Land Use and Land Cover change

This dataset contains a corpus of annotated text segments (sentences) extracted from scient