Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A consolidated lexical dataset for Dogon languages

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Hantgan, AbbieKpoSag
Éditeur:
Zenodo
Hôte:avatar

This dataset is a staged release of a consolidated lexical dataset for Dogon languages. It brings together heterogeneous source layers, including RefLex-derived data, Dogon and Bangime Linguistics materials, CLDF/LexiBank-derived working files, and subsequent BANG project curation. The workflow includes transcription standardization, source and language-name normalization, staged merging, Concepticon and part-of-speech enrichment, manual revision, source-village and GPS verification, doculect construction, and Glottolog alignment.

This release should be treated as provisional. Remaining issues include duplicate resolution, language-level attribution auditing, and verification of the incorporation of earlier manual curation of verbal paradigms. Full source-level attribution and contributor roles are documented in ATTRIBUTION.md.

Visit

doi.org

Languages

BangimeDogon, Yanda Dom

Tags

Language ScienceLexical SpreadsheetsLinguistic ComparisonLanguage DocumentationWest Africa

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Dogon languagesL-ReLF: A Framework for Lexical Dataset CreationA repository of free lexical resources for African languagesDataset: A consolidated and harmonised Verbal Autopsy dataset from Health and Demographic Surveillance Sites in South AfricaZA_LEX: lexical resources for South African languagesFikira Dataset | A Multilingual Reasoning Dataset for African Languages

Dogon languages

Media files collected in the Dogon languages and Bangime project (http://dogonlanguages.org).

L-ReLF: A Framework for Lexical Dataset Creation

This paper introduces the L-ReLF (Low-Resource Lexical Framework), a novel, reproducible methodology

A repository of free lexical resources for African languages

Dataset: A consolidated and harmonised Verbal Autopsy dataset from Health and Demographic Surveillance Sites in South Africa

This data note provides details of the development of a Verbal Autopsy (VA) dataset produced with th

ZA_LEX: lexical resources for South African languages

This repository contains lexical pronunciation resources and modules for use in text-to-speech (TTS) systems. Specifically, it was originally set up to track work on updating and enhancing existing resources for the NTTS project funded by the Department of Arts an

Fikira Dataset | A Multilingual Reasoning Dataset for African Languages

Fikira (Swahili for "thinking/reasoning") is a multilingual reasoning dataset for African languages, developed by Vambo AI. This dataset contains 50,000 reasoning examples across 10 African languages, synthetically generated as part of ongoing experiments at Vambo