Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

That doesn’t sound right: Evaluating speech transcription quality in field linguistics corpora

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssHarJiaLe
Éditeur:
Und
Hôte:avatar
Incorporating automatic speech recognition (ASR) into field linguistics workflows for language documentation has become increasingly common. While ASR performance has seen improvements in low-resource settings, obstacles remain when training models on data collected by documentary linguists. One notable challenge lies in the way that this data is curated. ASR datasets built from spontaneous speech are typically recorded in consistent settings and transcribed by native speakers following a set of well designed guidelines. In contrast, field linguists collect data in whatever format it is delivered by their language consultants and transcribe it as best they can given their language skills and the quality of the recording. This approach to data curation, while valuable for linguistic research, does not always align with the standards required for training robust ASR models. In this paper, we explore methods for identifying speech transcriptions in fieldwork data that may be unsuitable for training ASR models. We focus on two complimentary automated measures of transcription quality that can be used to identify transcripts with characteristics that are common in field data but could be detrimental to ASR training. We show that one of the metrics is highly effective at retrieving these types of transcriptions. Additionally, we find that filtering datasets using this metric of transcription quality reduces WER both in controlled experiments using simulated fieldwork with artificially corrupted data and in real fieldwork corpora.

Visit

doi.orgunderline.io

Tasks

automatic speech recognitionspeech processing

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similaires

Speech transcription platform speech servicesSpeech transcription serverWolof Speech TranscriptionSepedi Speech CorporaAfar language speech transcriptionOpenSLR African Speech Corpora

Speech transcription platform speech services

This is the Language Technology Services component implemented for the Speech Transcription Platform

Speech transcription server

This is the "Parliament-specific" application server component implemented as a proof-of-concept dur

Wolof Speech Transcription

Dataset de reconnaissance automatique de la parole (ASR) en wolof, une langue d'Afrique de l'Ouest p

Sepedi Speech Corpora

A corpus of Sesotho sa Leboa telephone speech data collected from mother tongue speakers of the sta

Afar language speech transcription

This dataset is designed to support the development of text-to-speech (TTS) and speech-to-text (STT)

OpenSLR African Speech Corpora

ASR system training, speech synthesis Notes / challenges: Includes crowd-sourced recordings for Yor