Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Innovative technologies for under-resourced language documentation: The BULB Project

Domaine:

natural language processing

Type de record:

paperproject
Créateur:
AddAddAmbBes
Éditeur:
LabLPPLanGro
Éditeur:
CCSD
Hôte:avatar
International audience The project Breaking the Unwritten Language Barrier (BULB), which brings together linguists and computer scientists, aims at supporting linguists in documenting unwritten languages. In order to achieve this we will develop tools tailored to the needs of documentary linguists by building upon technology and expertise from the area of natural language processing, most prominently automatic speech recognition and machine translation. As a development and test bed for this we have chosen three less-resourced African languages from the Bantu family: Basaa, Myene and Embosi. Work within the project is divided into three main steps: 1) Collection of a large corpus of speech (100h per language) at a reasonable cost. After initial recording, the data is re-spoken by a reference speaker to enhance the signal quality and orally translated into French. 2) Automatic transcription of the Bantu languages at phoneme level and the French translation at word level. The recognized Bantu phonemes and French words will then be automatically aligned. 3) Tool development. In close cooperation and discussion with the linguists, the speech and language technologists will design and implement tools that will support the linguists in their work, taking into account the linguists' needs and technology's capabilities. The data collection has begun for the three languages. For this we use standard mobile devices and a dedicated software—LIG-AIKUMA, which proposes a range of different speech collection modes (recording, respeaking, translation and elicitation). LIG-AIKUMA 's improved features include a smart generation and handling of speaker metadata as well as respeaking and parallel audio data mapping.

Visit

hal.science

Tasks

automatic speech recognitionmachine translationspeech processing

Languages

BasaaMbosiMyene

Tags

automatic alignmentunwritten languagesautomatic phonetic transcriptionLanguage documentation[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similaires

Breaking the unwritten language barrier: the BULB projectDocument Classification for the Under-resourced Amharic LanguageAuthor identification for Under-Resourced language (KadazanDusun)ibrahimbukhari1998/-Zero-Shot-for-Under-Resourced-LanguageShort Text Language Identification for Under Resourced LanguagesAutomatic speech recognition for an under-resourced language - amharic

Breaking the unwritten language barrier: the BULB project

International audience The project Breaking the Unwritten Language Barrier (BULB), wh

Document Classification for the Under-resourced Amharic Language

NLP is severely hampered by a scarcity of digital resources. This is especially true for Amharic, a

Author identification for Under-Resourced language (KadazanDusun)

This paper presents the task of Author Identification for KadazanDusun language by using

ibrahimbukhari1998/-Zero-Shot-for-Under-Resourced-Language

Zero-Shot POS Tagging for Under-Resourced Languages # Cross-Lingual POS Tagging: XLM-R vs. Glot500

Short Text Language Identification for Under Resourced Languages

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text languag

Automatic speech recognition for an under-resourced language - amharic