Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

RIYE Textual and Geospatial Corpus: An Aggregated Multidialectal Dataset for Low-Resource Language Processing

Domaine:

natural language processing

Type de record:

dataset
Créateur:
FolFolPopTho
Éditeur:
Fed
Éditeur:
Men
Hôte:avatar
This dataset integrates textual and spatial metadata collected under the Digiculture RIYE Project framework to support research in computational linguistics, spatial humanities, and low-resource dialect analysis. The corpus comprises three complementary datasets: RIYE_CSC-HTM_Aggregate.csv, which aggregates raw, qualitative text entries across cultural categories (such as food, festivals, music, history, and folklore) collected by the Computer Science (CSC) and Hospitality and Tourism Management (HTM) groups from the Federal University of Agriculture, Abeokuta; RIYE_CSC_Dataset.csv, which provides a structured matrix of thematic cultural elements (including dialect, clothing, religion, leadership, and performance arts) mapped across Local Government Areas (LGAs), towns/wards, and GIS location coordinates; and ogun_state_lgas_wards_coordinates.csv containing constitutionally recognised LGAs, wards, and coordinates extracted from the publicly available database compiled by Independent National Electoral Commission (INEC) before the 2023 General Elections. By combining unstructured descriptive text with spatially anchored domain attributes across regional zones, this collection establishes a rich baseline for requirement analysis, computational text processing, and geospatial modelling of dialectal and cultural elements.

Visit

doi.org

Tags

LinguisticsComputer ScienceArtificial IntelligenceRequirement EngineeringNatural Language ProcessingGIS DatabaseSpatial AnalysisTextual AnalysisApplied Machine Learning

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcode

Similaires

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

This dataset consists of a curated collection of high-fidelity, field-recorded audio samples develop