Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Spoken Term Detection Methods for Sparse Transcription in Very Low-resource Settings

Domain:

natural language processing

Record type:

paper
Creator:
FerBirBes
Host:avatar
We investigate the efficiency of two very different spoken term detection approaches for transcription when the available data is insufficient to train a robust ASR system. This work is grounded in very low-resource language documentation scenario where only few minutes of recording have been transcribed for a given language so far.Experiments on two oral languages show that a pretrained universal phone recognizer, fine-tuned with only a few minutes of target language speech, can be used for spoken term detection with a better overall performance than a dynamic time warping approach. In addition, we show that representing phoneme recognition ambiguity in a graph structure can further boost the recall while maintaining high precision in the low resource spoken term detection task.

Visit

arxiv.org

Tags

Computation and LanguageSoundAudio and Speech Processing

Similar

Word-based Probabilistic Phonetic Retrieval for Low-resource Spoken Term DetectionVoice Conversion Can Improve ASR in Very Low-Resource SettingsUsing Machine Learning for Medical Error Detection in Low-Resource SettingsEvansKonadu/AI-Powered-Pneumonia-Detection-in-Low-Resource-SettingsA Multi-Task Benchmark for Abusive Language Detection in Low-Resource SettingsFinetuning Pretrained Model with Embedding of Domain and Language Information for ASR of Very Low-Resource Settings

Word-based Probabilistic Phonetic Retrieval for Low-resource Spoken Term Detection

Two problems make Spoken Term Detection (STD) particularly challenging under low-resource conditi

Voice Conversion Can Improve ASR in Very Low-Resource Settings

Voice conversion (VC) could be used to improve speech recognition systems in low-resource languages

Using Machine Learning for Medical Error Detection in Low-Resource Settings

Abstract Medication errors during surgical procedures pose significant risks to pa

EvansKonadu/AI-Powered-Pneumonia-Detection-in-Low-Resource-Settings

Deep learning models for pneumonia detection in chest X-rays, optimised for deployment in resource-c

A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings

Content moderation research has recently made significant advances, but remains limited in serving t

Finetuning Pretrained Model with Embedding of Domain and Language Information for ASR of Very Low-Resource Settings

This study investigates the effective incorporation of meta-information such as domain and language