Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

WALA: A Multilingual Resource Repository for West African Languages

Domain:

natural language processing

Record type:

paper
The West African Language Archive (WALA) initiative has emerged from a number of concurrent projects, and aims to encourage local scholars to create high quality decentralised repositories documenting West African languages, and to make these repositories available to language communities, language planners, educationalists and scientists via an internet metadata portal such as OLAC (Open Language Archive Community). A wide range of criteria has to be met in designing and implementing this kind of archive. We discuss these criteria with reference to experiences in documentation work in three very different ongoing language documentation projects, on designing an encyclopaedia, on documenting an endangered language, and on creating a speech synthesiser. We pay special attention to the provision of metadata, a formal variety of catalogue or housekeeping information, without which resources are doomed to remain inaccessible.

Visit

aclanthology.org

Tags

acl

Similar

GrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph KnowledgeAfrican Voices: Multilingual Speech Dataset for Low-Resource African LanguagesA repository of free lexical resources for African languagesTriLex: A Framework for Multilingual Sentiment Analysis in Low-Resource South African LanguagesWebCrawl African : A Multilingual Parallel Corpora for African LanguagesMultilingual Neural Machine Translation for Zero-Resource Languages

GrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph Knowledge

Contextualized embeddings based on large language models (LLMs) are available for various languages,

African Voices: Multilingual Speech Dataset for Low-Resource African Languages

A large-scale multilingual speech dataset developed by Data Science Nigeria. Contains more than 3,000 hours of transcribed audio across four Nigerian languages: Hausa, Igbo, Nigerian Pidgin, and Yorùbá. The dataset supports Automatic Speech Recognition (ASR) and sp

A repository of free lexical resources for African languages

TriLex: A Framework for Multilingual Sentiment Analysis in Low-Resource South African Languages

Low-resource African languages remain underrepresented in sentiment analysis, limiting both lexical

WebCrawl African : A Multilingual Parallel Corpora for African Languages

WebCrawl African is a mixed domain multilingual parallel corpora for a pool of African languages com

Multilingual Neural Machine Translation for Zero-Resource Languages

In recent years, Neural Machine Translation (NMT) has been shown to be more effective than phrase-ba