Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

karya-inc/dataset-mundari-tts

Domaine:

natural language processing

Type de record:

dataset
Créateur:
kar
Hôte:
# Mundari TTS Dataset Corpus This dataset contains 26,870 recordings of Mundari speech, spoken by two different speakers (one female and one male). It is released under the non-commercial version of the Karya Public License. Please read the "License" section below for a high-level summary of what you are allowed and not allowed to do under this license. This corpus was created by Microsoft Research India, the Indian Institute of Technology Kharagpur, and Karya in collaboration with the project of Indo-German Development Cooperation “FAIR Forward – Artificial Intelligence for all”. FAIR Forward is being implemented by Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ) on behalf of the Federal Ministry for Economic Cooperation and Development (BMZ). ## Data Description The Mundari TTS dataset corpus contains a total of 26,870 audio files, each containing a single utterance spoken by one of the two speakers. The audio is recorded in 32-bit PCM format with a sampling rate of 44.1 kHz. The dataset includes transcripts for each audio file in Mundari script. The recordings were collected in a sound-treated room using a high-quality microphone and preamp. ## Speakers The Mundari TTS dataset corpus includes recordings from two different speakers: a female speaker and a male speaker. The female speaker contributed 19,868 recordings, while the male speaker contributed 7,002 recordings. The total size of the corpus is approximately 17 GB (7 GB compressed). Therefore, the full dataset is hosted in cloud storage (instead of this github repository). This repository contains a sample of 100 recordings from each speaker. Please email data@karya.in for a link to download the full dataset. ## Contents The repository contains the following files: 1. `README.md` - This file. 2. `LICENSE.txt` - The full text of the license under with this dataset is released. 3. `COPYRIGHT.txt` - Copyright for the dataset. 4. `data-sample.tgz` - Sample data folder. It contains two …

Visit

github.com

Tasks

speech processingtext to speech

Languages

Kirya-KonzelMandari

Similaires

karya-inc/parikshaMundari-English dictionary Mundari dictionaryMbosi-TTS-DatasetRw Tts DatasetYaka-TTS-DatasetEton-TTS-Dataset

karya-inc/pariksha

This repository contains the data generated by Karya for Pariksha. More details about this project c

Mundari-English dictionary Mundari dictionary

https://www.sil.org/resources/archives/55041

Mbosi-TTS-Dataset

The dataset consists of paired audio and text data on Mbosi (mdw), a language spoken in Congo. The a

Rw Tts Dataset

Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, co

Yaka-TTS-Dataset

Paired audio and text data on Yaka (also known as West Teke), a language spoken in Congo. The audio

Eton-TTS-Dataset

Eton-TTS-Dataset is a single-speaker scripted speech dataset dedicated to the documentation and tech