Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

About vocabulary adaptation for automatic speech recognition of video data

Domain:

natural language processing

Record type:

paper
Creator:
JouLanMenFoh
Editor:
SpeStaGri
Publisher:
CCSD
Host:avatar
International audience This paper discusses the adaptation of vocabularies for automatic speech recognition. The context is the transcriptions of videos in French, English and Arabic. Baseline automatic speech recognition systems have been developed using available data. However, the available text data, including the GigaWord corpora from LDC, are getting quite old with respect to recent videos that are to be transcribed. The paper presents the collection of recent textual data from internet for updating the speech recognition vocabularies and training the language models, as well as the elaboration of development data sets necessary for the vocabulary selection process. The paper also compares the coverage of the training data collected from internet, and of the GigaWord data, with finite size vocabularies made of the most frequent words. Finally, the paper presents and discusses the amount of out-of-vocabulary word occurrences, before and after update of the vocabularies, for the three languages.

Visit

inria.hal.science

Tasks

automatic speech recognitionlanguage modelingspeech processing

Tags

vocabulary selectionvocabulary adaptationvocabularySpeech recognition[INFO.INFO-TS]Computer Science [cs]/Signal and Image Processing

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similar

Generative Adversarial Training Data Adaptation for Very Low-resource Automatic Speech RecognitionAutomatic Speech Recognition And Limited Vocabulary: A SurveyDevelopment of a diacritic-aware large vocabulary automatic speech recognition for Hausa languageA Multi-Faceted Framework for Personalizing Automatic Speech Recognition for Impaired Speech: Combining Bayesian Adaptation and Semantic Data SynthesisAwezaMed automatic speech recognition (ASR) test dataSynthetic Voice Data for Automatic Speech Recognition in African Languages

Generative Adversarial Training Data Adaptation for Very Low-resource Automatic Speech Recognition

It is important to transcribe and archive speech data of endangered languages for preserving heritag

Automatic Speech Recognition And Limited Vocabulary: A Survey

Automatic Speech Recognition (ASR) is an active field of research due to its large number of applica

Development of a diacritic-aware large vocabulary automatic speech recognition for Hausa language

A Multi-Faceted Framework for Personalizing Automatic Speech Recognition for Impaired Speech: Combining Bayesian Adaptation and Semantic Data Synthesis

Automatic Speech Recognition (ASR) systems consistently fail for individuals with non-normative spee

AwezaMed automatic speech recognition (ASR) test data

The corpus contains orthographically transcribed broadband speech in four official languages of So

Synthetic Voice Data for Automatic Speech Recognition in African Languages

Speech technology remains out of reach for most of the over 2300 languages in Africa. We present the first systematic assessment of large-scale synthetic voice corpora for African ASR.

We apply a three-step process: LLM-driven text creation, TTS voice sy