Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Operation LiLi: Using Crowd-Sourced Data and Automatic Alignment to Investigate the Phonetics and Phonology of Less-Resourced Languages

Domaine:

natural language processing

Type de record:

paperdatasetsoftware
Créateur:
Hutin, MathildeAll
Éditeur:
LabTraÉcoANR
Éditeur:
CCSDMDPI
Hôte:avatar
International audience Less-resourced languages are usually left out of phonetic studies based on large corpora. We contribute to the recent efforts to fill this gap by assessing how to use open-access, crowd-sourced audio data from Lingua Libre for phonetic research. Lingua Libre is a participative linguistic library developed by Wikimedia France in 2015. It contains more than 670k recordings in approximately 150 languages across nearly 740 speakers. As a proof of concept, we consider the Inventory Size Hypothesis, which predicts that, in a given system, variation in the realization of each vowel will be inversely related to the number of vowel categories. We investigate data from 10 languages with various numbers of vowel categories, i.e., German, Afrikaans, French, Catalan, Italian, Romanian, Polish, Russian, Spanish, and Basque. Audio files are extracted from Lingua Libre to be aligned and segmented using the Munich Automatic Segmentation System. Information on the formants of the vowel segments is then extracted to measure how vowels expand in the acoustic space and whether this is correlated with the number of vowel categories in the language. The results provide valuable insight into the question of vowel dispersion and demonstrate the wealth of information that crowd-sourced data has to offer.

Visit

hal.science

Tasks

speech processing

Languages

Afrikaans

Tags

[SCCO.LING]Cognitive science/Linguistics

Licenses

info:eu-repo/semantics/OpenAccess

Similaires

End-To-End Multilingual Automatic Speech Recognition For Less-Resourced Languages: The Case Of Four Ethiopian LanguagesA Comparative Study of the Phonetics and Phonology of Surmic LanguagesAn Evaluation of Spatial Network Modeling To Aid Sanitation Planning In Informal Settlements Using Crowd-Sourced DataPhonology and PhoneticsStudies in Kasem phonetics and phonology Kasem phonetics and phonology Kasem studiesThe phonetics and phonology of geminate consonants

End-To-End Multilingual Automatic Speech Recognition For Less-Resourced Languages: The Case Of Four Ethiopian Languages

Presenter: Solomon Teferra Abate, Martha Yifiru Tachbelie, Tanja Schultz , ICASSP 20

A Comparative Study of the Phonetics and Phonology of Surmic Languages

An Evaluation of Spatial Network Modeling To Aid Sanitation Planning In Informal Settlements Using Crowd-Sourced Data

Abstract: Limited water and sanitation infrastructure in rapidly urbanising informal settlements can

Phonology and Phonetics

Abstract This chapter focuses on the contributions African languages have made to phonological theo

Studies in Kasem phonetics and phonology Kasem phonetics and phonology Kasem studies

Includes bibliographical references (p. [131]-136)

The phonetics and phonology of geminate consonants

International audience This talk deals with the phonetic implementation and the phono