Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

RIYE Textual and Geospatial Corpus: An Aggregated Multidialectal Dataset for Low-Resource Language Processing

Domain:

natural language processing

Record type:

dataset
Creator:
FolFolPopTho
Editor:
Fed
Publisher:
Men
Host:avatar
This dataset integrates textual and spatial metadata collected under the Digiculture RIYE Project framework to support research in computational linguistics, spatial humanities, and low-resource dialect analysis. The corpus comprises three complementary datasets: RIYE_CSC-HTM_Aggregate.csv, which aggregates raw, qualitative text entries across cultural categories (such as food, festivals, music, history, and folklore) collected by the Computer Science (CSC) and Hospitality and Tourism Management (HTM) groups from the Federal University of Agriculture, Abeokuta; RIYE_CSC_Dataset.csv, which provides a structured matrix of thematic cultural elements (including dialect, clothing, religion, leadership, and performance arts) mapped across Local Government Areas (LGAs), towns/wards, and GIS location coordinates; and ogun_state_lgas_wards_coordinates.csv containing constitutionally recognised LGAs, wards, and coordinates extracted from the publicly available database compiled by Independent National Electoral Commission (INEC) before the 2023 General Elections. By combining unstructured descriptive text with spatially anchored domain attributes across regional zones, this collection establishes a rich baseline for requirement analysis, computational text processing, and geospatial modelling of dialectal and cultural elements.

Visit

doi.org

Tags

LinguisticsComputer ScienceArtificial IntelligenceRequirement EngineeringNatural Language ProcessingGIS DatabaseSpatial AnalysisTextual AnalysisApplied Machine Learning

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcode

Similar

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

This dataset consists of a curated collection of high-fidelity, field-recorded audio samples develop