Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus

Domain:

natural language processing

Record type:

paperdataset
Creator:
Ouz
Host:avatar
Aligned audio corpora are fundamental to NLP technologies such as ASR and speech translation, yet they remain scarce for underrepresented languages, hindering their technological integration. This paper introduces a methodology for constructing LoReSpeech, a low-resource speech-to-speech translation corpus. Our approach begins with LoReASR, a sub-corpus of short audios aligned with their transcriptions, created through a collaborative platform. Building on LoReASR, long-form audio recordings, such as biblical texts, are aligned using tools like the MFA. LoReSpeech delivers both intra- and inter-language alignments, enabling advancements in multilingual ASR systems, direct speech-to-speech translation models, and linguistic preservation efforts, while fostering digital inclusivity. This work is conducted within Tutlayt AI project (tutlayt.fr). This paper is withdrawn because the LoReSpeech dataset described in Section 2 is not currently available, which affects the reproducibility of the work and the validity of the experimental results

Visit

arxiv.org

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Tags

Computation and Language

Similar

IWSLT2022 - Low-resource Speech Translation Track: Tamasheq-French Parallel CorpusJW300: A Wide-Coverage Parallel Corpus for Low-Resource LanguagesEthioMT: Parallel Corpus for Low-resource Ethiopian LanguagesA Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation BenchmarksFiltered Pseudo-parallel Corpus Improves Low-resource Neural Machine TranslationAfrican Voices: Multilingual Speech Dataset for Low-Resource African Languages

IWSLT2022 - Low-resource Speech Translation Track: Tamasheq-French Parallel Corpus

Repository for sharing the data in the Tamasheq language, one of the languages for the low-resource speech translation track at IWSLT 2022.

JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages

EthioMT: Parallel Corpus for Low-resource Ethiopian Languages

Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT leaves much to be desired for

A Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation Benchmarks

Filtered Pseudo-parallel Corpus Improves Low-resource Neural Machine Translation

Large-scale parallel corpora are essential for training high-quality machine translation systems; ho

African Voices: Multilingual Speech Dataset for Low-Resource African Languages

A large-scale multilingual speech dataset developed by Data Science Nigeria. Contains more than 3,000 hours of transcribed audio across four Nigerian languages: Hausa, Igbo, Nigerian Pidgin, and Yorùbá. The dataset supports Automatic Speech Recognition (ASR) and sp