Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Santali Emotional Speech Corpus)

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Das
Éditeur:
Zenodo
Hôte:avatar

Specifications Table

Subject

Signal Processing

Specific Subject Area

Speech Emotion Recognition, Sound Analysis

Type of Data

Digital Audio Files

How the Data were acquired

Santali language voice recordings are done in a noiseless environment using a laptop, external microphone and headset. The recordings, initially in MP3 format, are converted to WAV using FFmpeg. The dataset comprises approximately 1 hour and 53 minutes of audio, which is cut into 3–5 second duration samples. Background noise removal and segmentation are carried out using the Audacity software.

Tools used are:

·      Microphone

·      Headset

·      Laptop

·      FFmpeg Software

·      Audacity Software

Data Format

Waveform Audio File Format (WAV)

Description of Data Collection

The Speech have been recorded in different categories:

·      Five Different Emotional States (Angry, Happy, Neutral, Sad, Surprise).

·      Using Five Different Santali Sentences:

                 i.      Santali, Ol Chiki: ᱤᱭᱮ ᱪᱟᱱᱫᱚ ᱟᱢ ᱮᱠᱞᱟ ᱛᱮᱠᱮ ᱵᱤᱞ ᱮᱢᱟ

English: You will have to pay this month’s bill completely by yourself.

                ii.      Santali, Ol Chiki: ᱤᱭᱮ ᱪᱟᱱᱫᱚ ᱵᱟᱛᱭᱟ ᱫᱟᱱ ᱦᱮᱡ ᱫᱟᱨᱮᱭᱟ

English: This month, Cyclone Dana may come.

              iii.      Santali, Ol Chiki: ᱪᱟᱞᱟ ᱠᱟᱴᱟᱠ ᱨᱮ ᱱᱟ ᱫᱟᱦᱤᱵᱟᱨᱟ ᱢᱮᱱᱠᱼ ᱡᱟᱢ

English: Let’s go to Cuttack and eat Dahibara.

               iv.      Santali, Ol Chiki: ᱤᱭᱮ ᱨᱟᱵᱤᱵᱟᱨ ᱟᱵᱚ ᱱᱟᱱᱫᱟᱱᱠᱟᱱᱟᱱ ᱪᱟᱞᱟ ᱵᱮ ᱟᱵᱚ ᱱᱟ

English: This Sunday, we have to go to Nandankanan.

                v.      Santali, Ol Chiki: ᱤᱭᱮ ᱨᱟᱵᱤᱵᱟᱨ ᱦᱟᱴ ᱨᱮ ᱟᱞᱩ ᱫᱟᱨ ᱠᱳᱰᱤᱮ ᱴᱟᱱᱠᱟ ᱠᱤᱞᱚ ᱟᱨ ᱯᱤᱭᱟᱡ ᱫᱟᱨ ᱛᱤᱨᱤᱥ ᱴᱟᱱᱠᱟ ᱠᱤᱞᱚ

English: This Sunday at the market, potatoes cost twenty rupees per kilo and onions cost thirty rupees per kilo.

A total of 2000 Santali speech recordings were collected from 80 speakers of diverse backgrounds across different regions of Koraput, Odisha, within the age group of 18 to 60 years, where the emotions were simulated and enacted. The original five statements, initially in Odia, were translated into Santali (Ol Chiki script) by a proficient bilingual speaker, and participants were given detailed instructions regarding the scripted content and intended emotional states prior to recording. The dataset is class-balanced, with each of the five emotions containing 400 samples and each statement contributing 80 samples per emotion. All recordings were sampled at 48 kHz, resulting in a total dataset size of approximately 1.46 GB. The emotional labels were further validated by ten human evaluators, achieving an average recognition accuracy of about 74.6% for the intended emotions in the proposed SantaliESC Speech-Emotion Dataset.

Data Source Location

City: Koraput

State: Odisha

Country: India

Data Accessibility

Repository Name: SantaliESC: A Santali Emotional Speech Corpus.

Digital Object Identifier: 10.5281/zenodo.20702265

Funding

This work has been supported by the Telecom Centre of Excellence India under the Telecom Technology Development Fund (TTDF).

Visit

doi.org

Tasks

emotion identificationspeech processing

Languages

Ndasa

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

AlbEmo: a perceptually validated emotional speech corpus for AlbanianA PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATIONWhispering in Ol Chiki: Cross-Lingual Transfer Learning for Santali Speech RecognitionArabic Speech CorpusNCHLT speech corpusHausa Speech Corpus

AlbEmo: a perceptually validated emotional speech corpus for Albanian

AlbEmo is, to our knowledge, the first emotional speech corpus for Albanian, a language not covered

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

Machine Translation (MT) poses a significant challenge in developing language corpora for low-resour

Whispering in Ol Chiki: Cross-Lingual Transfer Learning for Santali Speech Recognition

Arabic Speech Corpus

This Speech corpus has been developed as part of PhD work carried out by Nawar Halabi at the Univers

NCHLT speech corpus

The NCHLT speech corpus of the South African languages

Hausa Speech Corpus

This is a Hausa Speech data set that was recorded as a baseline for Hausa Speech Recognition. The data sets can be used in building Automatic Speech recognition for Hausa language, Speech synthesis and speaker recognition.