Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

terrillschrock/east-african-lexome

Domaine:

natural language processing

Type de record:

dataset
Créateur:
ter
Hôte:
A growing collection of lexical datasets for East African languages # East African Lexome (EALex) A comprehensive, machine-readable comparative lexical database for the languages of East Africa. EALex is research infrastructure for African historical linguistics in the era of computational and quantitative methods. The project aims to bring "lexomics" — the science of large-scale historical lexical data — to a region whose linguistic richness has long been studied through patchwork sources rather than as an integrated whole. ## Scope EALex covers **420 languages** spoken across **ten East African nations**: Burundi, Eritrea, Ethiopia, Kenya, Rwanda, Somalia, South Sudan, Sudan, Tanzania, and Uganda. The languoid inventory is built on Glottolog 5.2 with project-specific reconciliation decisions (documented per-language in the dataset) where Ethnologue, regional expertise, or geographic considerations warranted divergence from Glottolog's classification. The languages span four major language families plus several isolates: - **Atlantic-Congo** (largely Bantu) — 183 languages - **Nilo-Saharan** (Nilotic, Central Sudanic, Surmic, Koman, Kordofanian, and others) — combined majority - **Afro-Asiatic** (Cushitic, Omotic, Ethiosemitic, Berber-Beja) — 67 languages - **Isolates** — 10 languages, including Hadza, Sandawe, Kunama, Nara, Berta, Ongota, and Shabo Over half of the languages in scope are classified as threatened, shifting, moribund, or nearly extinct under Glottolog's Agglomerated Endangerment Status (AES) ratings — making documentation efforts time-sensitive. ## Goal The project's working target is **one million attested lexical forms** — an average of roughly 2,381 words per language, though some languages will contribute many more (an established dictionary like the project lead's own Ik dictionary contributes around 4,000 verb and noun roots alone) and others far fewer. ## Data format EALex is published in CLDF (Cross-Linguistic Data Format), the standard format for comparative linguistic datasets. CLDF makes the da …

Visit

github.com

Languages

BerberHadzaKunamaOmotikOngotaSandaweShabo

Tags

africadictionarieslexicographylexomicslinguisticswordlists

Similaires

East African EnglishEast African NubiEast African CybersecurityEast African fossil recordEast African Religious Pluralismoduorojuang/East-African-Pronunciation-

East African English

English in East Africa is a well-developed usage variety (or a cluster of usage varieties), although

East African Nubi

SUMMARY A central question in Creole studies has been to ascertain to what degree the structure of

East African Cybersecurity

ICT has become an essential factor in Ethiopia and other East African countries to transform their e

East African fossil record

For almost a century East Africa has been a prime location for paleoanthropological research. Ethiop

East African Religious Pluralism

Abstract For many recent generations the city of Mombasa, Kenya, on the east African coast (pop. 1

oduorojuang/East-African-Pronunciation-

AugSt is a phonetic translation system that renders English Proper into readable, speakable phonetic