Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Towards a Crowdsourcing Platform for Low Resource Languages -- A Collectivist Approach

Domain:

natural language processing

Record type:

project
Creator:
AssHomLeventhal, MichaelLug
Publisher:
Und
Host:avatar
This work demonstrates how semi-supervised learning and human-in-the-loop crowdsourcing can help neural machine translation (NMT) challenges common in low-resource languages. We focus on the Mande language Bambara, which has approximately 16 million primary and secondary speakers in Western Africa. Bambara is mainly spoken as opposed to written language and it has few digital resources due to its history in regions where colonial French became the language of government and industry. Thus, Bambara is a "low-resource language" and because it lacks the existing language resources (parallel digital text and labeled data) necessary for NMT, we describe a novel crowdsourcing approach to support semi-supervised NMT. We designed a crowdsourcing platform that requests the annotator to supply information when the NMT model has decision confusion. Our crowdsourcing platform was tested on evaluating translations of Malian broadcast news and Wikipedia pages in Bambara. Our initial research shows a wide variation in the quality of the translations and further work includes a more rigorous evaluation of translator skills when onboarding new annotators.

Visit

doi.orgunderline.io

Tasks

machine translation

Languages

BamanankanLameMandinkaManinkakan, EasternNomaandeTobanga

Tags

Labour EconomicsComputational IntelligenceHuman-Computer Interaction

Similar

Reusable Component Retrieval: A Semantic Search Approach for Low-Resource LanguagesIndiAnn: An Annotation Platform for Low-Resource Indic LanguagesTowards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse LanguagesTowards Guided Back-translation for Low-resource languages- A Case Study on Kabyle-FrenchThe Philotis Platform: Empowering Low-Resource Languages ProcessingDetecting Urgency Status of Crisis Tweets: A Transfer Learning Approach for Low Resource Languages

Reusable Component Retrieval: A Semantic Search Approach for Low-Resource Languages

A common practice among programmers is to reuse existing code, accomplished by performing natural la

IndiAnn: An Annotation Platform for Low-Resource Indic Languages

Linguistic annotation tools that work well for non-Indic languages (e.g. English, German, Spanish, e

Towards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse Languages

Parallel corpora are ideal for extracting a multilingual named entity (MNE) resource, i.e., a datase

Towards Guided Back-translation for Low-resource languages- A Case Study on Kabyle-French

The Philotis Platform: Empowering Low-Resource Languages Processing

The project Philotis has developed a web-based platform1 and a methodology, implemented as a pipelin

Detecting Urgency Status of Crisis Tweets: A Transfer Learning Approach for Low Resource Languages

We release an urgency dataset that consists of English tweets relating to natural crises. The set is