Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
Buzaaba, HappyDioAdeKah
Host:avatar
Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntactically annotated treebanks for nine diverse African languages spanning major language families and regions across Sub-Saharan Africa. Using the Surface-Syntactic Universal Dependencies (SUD) framework, our community-led effort provides high-quality, native-speaker verified data that capture typological key features such as agglutination and tone. We evaluate a range of models on AfriSUD for part-of-speech tagging and dependency parsing including non-transformer baselines, multilingual pretrained encoders, and LLMs. Our results reveal a significant syntax gap, where models still show clear limitations across the nine languages, suggesting that existing architectures may not fully capture the structural diversity of African-language syntax.

Visit

arxiv.org

Tasks

dependency parsingparsingpart of speech tagging

Tags

Computation and LanguageArtificial Intelligence

Similar

Afribooms Afrikaans Dependency TreebankOludeji/Yoruba-Dependency-TreebankThe First Universal Dependency Treebank for Tswana: Tswana-PopapoleloDependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource LanguagesCross-Dialectal Transfer for Low-Resource Arabic: The Tunisian Arabic Dependency TreebankCIRAL: A Test Collection for CLIR Evaluations in African Languages

Afribooms Afrikaans Dependency Treebank

This is the annotated corpus developed for Afrikaans for the Afribooms project. The corpus includes

Oludeji/Yoruba-Dependency-Treebank

A Universal Dependencies (UD) treebank for the Yorùbá language, annotated for linguistic and computa

The First Universal Dependency Treebank for Tswana: Tswana-Popapolelo

Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages

Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, ye

Cross-Dialectal Transfer for Low-Resource Arabic: The Tunisian Arabic Dependency Treebank

CIRAL: A Test Collection for CLIR Evaluations in African Languages