Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Burushaski-English Speech Translation Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Hab
Host:
The Burushaski Speech–English Parallel Corpus is a community-driven language resource developed to support research in speech translation, machine translation, automatic speech recognition, and other language technologies for Burushaski, an under-resourced language isolate spoken in northern Pakistan. The dataset was collected using an audio-first, linguistically informed framework that combines structured elicitation targeting high-frequency vocabulary and key grammatical phenomena with the collection of functional and conversational language relevant to real-world communication. Data collection was facilitated through a custom mobile application that standardized prompts while enabling scalable community participation, and the development process incorporated continuous feedback from Burushaski-speaking contributors to improve linguistic coverage and cultural relevance. The current pilot release contains 14,970 recorded utterances from native speakers, with each audio recording paired with an English translation, creating a parallel corpus intended to advance research, promote language inclusion in AI, and expand the digital presence of Burushaski.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Tags

mdcmozilla data collectiveMTWAVTXT

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)

Similar

BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation CorpusAfrican Speech Technology English-English Speech CorpusMaasai-English Translation CorpusA Corpus for Amharic-English Speech Translation: The Case of Tourism DomainNCHLT English Speech CorpusTeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation

BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus

There is a major shortage of Speech-to-Speech Translation (S2ST) datasets for high resource-to-low r

African Speech Technology English-English Speech Corpus

African Speech Technology speech and transcription data for the English-English database. The "spee

Maasai-English Translation Corpus

Parallel English↔Maasai translation pairs for low-resource MT, language preservation, and culturally

A Corpus for Amharic-English Speech Translation: The Case of Tourism Domain

Speech translation research for the major languages like English, Japanese and Spanish has been conducted since the 1980’s. But no attempt were made in speech translation to/from the under-resourced language like Amharic. These activities suffered from the lack of

NCHLT English Speech Corpus

Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test sui

TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation

Despite Telugu being spoken by over 80 million people, speech translation research for this morpholo