Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Investigating Language Impact in Bilingual Approaches for Computational Language Documentation

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
BoiVilBes
Hôte:avatar
For endangered languages, data collection campaigns have to accommodate the challenge that many of them are from oral tradition, and producing transcriptions is costly. Therefore, it is fundamental to translate them into a widely spoken language to ensure interpretability of the recordings. In this paper we investigate how the choice of translation language affects the posterior documentation work and potential automatic approaches which will work on top of the produced bilingual corpus. For answering this question, we use the MaSS multilingual speech corpus (Boito et al., 2020) for creating 56 bilingual pairs that we apply to the task of low-resource unsupervised word segmentation and alignment. Our results highlight that the choice of language for translation influences the word segmentation performance, and that different lexicons are learned by using different aligned translations. Lastly, this paper proposes a hybrid approach for bilingual word segmentation, combining boundary clues extracted from a non-parametric Bayesian model (Goldwater et al., 2009a) with the attentional word segmentation neural model from Godard et al. (2018). Our results suggest that incorporating these clues into the neural models' input representation increases their translation and alignment quality, specially for challenging language pairs. Accepted to 1st Joint SLTU and CCURL Workshop

Visit

arxiv.org

Tasks

speech processing

Tags

Computation and Language

Similaires

Integrating descriptive and computational approaches in language documentation and resource developmentUnsupervised word discovery for computational language documentationComputational Approaches to Understanding Large Language Model Impact on Writing and Information EcosystemsA Very Low Resource Language Speech Corpus for Computational Language Documentation ExperimentsAfrican language documentation: new data, methods and approachesBULBasaa: A Bilingual Basaa-French Speech Corpus for the Evaluation of Language Documentation Tools

Integrating descriptive and computational approaches in language documentation and resource development

The benefits of interdisciplinary teams as well as the creation of documentary products of a variety

Unsupervised word discovery for computational language documentation

Découverte non-supervisée de mots pour outiller la linguistique de terrain La dive

Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems

Large language models (LLMs) have shown significant potential to change how we write, communicate, a

A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments

Most speech and language technologies are trained with massive amounts of speech and text informatio

African language documentation: new data, methods and approaches

National Foreign Language Resource Center Seyfeddinipur_2016.pdf

BULBasaa: A Bilingual Basaa-French Speech Corpus for the Evaluation of Language Documentation Tools