Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Tusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
MorPicLi,Sim
Hôte:avatar
There is growing interest in ASR systems that can recognize phones in a language-independent fashion. There is additionally interest in building language technologies for low-resource and endangered languages. However, there is a paucity of realistic data that can be used to test such systems and technologies. This paper presents a publicly available, phonetically transcribed corpus of 2255 utterances (words and short phrases) in the endangered Tangkhulic language East Tusom (no ISO 639-3 code), a Tibeto-Burman language variety spoken mostly in India. Because the dataset is transcribed in terms of phones, rather than phonemes, it is a better match for universal phone recognition systems than many larger (phonemically transcribed) datasets. This paper describes the dataset and the methodology used to produce it. It further presents basic benchmarks of state-of-the-art universal phone recognition systems on the dataset as baselines for future experiments. 4 pages, 3 figures

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similaires

An Empirical Recipe for Universal Phone RecognitionSagalee: an Open Source Automatic Speech Recognition Dataset for Oromo LanguageUniversal Phone Recognition with a Multilingual Allophone SystemVoxMg: An Automatic Speech Recognition Dataset for MalagasyEnabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan LanguageIn-context Language Learning for Endangered Languages in Speech Recognition

An Empirical Recipe for Universal Phone Recognition

Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, ye

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoke

Universal Phone Recognition with a Multilingual Allophone System

Multilingual models can improve language processing, particularly for low resource situations, by sh

VoxMg: An Automatic Speech Recognition Dataset for Malagasy

African languages are not well-represented in Natural Language Processing (NLP). The main reason is a lack of resources for training models. Low-resource languages, such as Malagasy, cannot benefit from modern NLP methods if no datasets are available. This paper pr

Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language

The lack of impaired speech data hinders advancements in the development of inclusive speech technol

In-context Language Learning for Endangered Languages in Speech Recognition

With approximately 7,000 languages spoken worldwide, current large language models (LLMs) support on