Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tujia speech features from educational videos: pilot dataset (110 utterances)

Domain:

natural language processing

Record type:

dataset
Creator:
XiaCuiXiaQiu
Publisher:
Zenodo
Host:avatar
This package shares the derived acoustic features and quality metrics for 110 Tujia utterances selected from the public educational video series "Gen Wo Xue Shuo Tujia Yu" (跟我学说土家语). It does not contain raw video, raw audio, screenshots, speaker images, or personally identifiable information. Raw media cannot be redistributed because of platform copyright restrictions. Processing summary: Audio was standardized to 16,000 Hz, mono, 16-bit PCM, peak amplitude 0.99. A first-order pre-emphasis filter H(z) = 1 - 0.97 z^-1 was applied. Spectral subtraction was used for noise reduction. Voice activity detection used short-time energy and zero-crossing rate (25 ms frames, 10 ms shift, Hamming window; noise floor from the lowest-energy frames; high/low energy thresholds of noise floor +8/+3 dB; ZCR threshold 0.25; gaps shorter than 150 ms merged; 200 ms margins). Quality metrics: speech frame ratio >= 30%, SNR >= 10 dB (capped at 60 dB), clipping ratio < 1%, spectral dynamic range reported descriptively, duration consistency 0.3-20 s. Feature extraction: F0 range 50-500 Hz (normalized cross-correlation); LPC order 18 with formant constraints of 150-5000 Hz and bandwidth < 400 Hz; MFCC with 13 coefficients, 26 Mel filters, and 2048-point FFT. Contents: acoustic feature data (F0, F1/F2, mean MFCC C1-C5), quality assessment metrics, processing parameters, data dictionary, and the MATLAB processing pipeline (scripts/). Related manuscript: "Acoustic analysis of Tujia speech from educational videos: A low-resource dataset study". Ethics and copyright: The source videos are public educational content posted for language teaching purposes. No raw audio/video, video URLs, or creator-identifying metadata are included. The data were collected from open-access platforms without direct interaction with human subjects.

Visit

doi.org

Tasks

speech processing

Tags

Tujia languagelow-resource speech datasetacoustic analysisfundamental frequencyformantsMFCCendangered languages

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Speech act analysis of Igbo utterances in funeral ritesA multi-site loiasis dataset of whole blood videosKenya News Article Features DatasetSIZ_VS_NARR-110 SiwiRcrops Luganda Utterances MFCCsActivity Recognition From Newborn Resuscitation Videos

Speech act analysis of Igbo utterances in funeral rites

A multi-site loiasis dataset of whole blood videos

Loiasis is a blood-borne filarial infection under consideration for inclusion in the WHO’s priority

Kenya News Article Features Dataset

Kenya News Article Features Dataset (Metadata + NLP + News Structure Signals)

SIZ_VS_NARR-110 Siwi

Explanation about different parts of Siwa Equipment: Edirol R-O5 recorder / Audiotechnica AT897 Mono

Rcrops Luganda Utterances MFCCs

Activity Recognition From Newborn Resuscitation Videos

Objective: Birth asphyxia is one of the leading causes of neonatal deaths. A key for survival is per