Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TCST-6D-AD: A Six-tuple Parallel Tibetan-Chinese Speech Translation Dataset for the Amdo Dialect TCST-6D-AD: 面向安多方言的藏汉六元组平行语音翻译数据集

Domain:

natural language processing

Record type:

dataset
Creator:
CuoCanDoraMao
Publisher:
Sci
Host:avatar
We constructed the TCST-6D-AD end-to-end Tibetan-Chinese speech translation dataset. The source data was obtained through three channels: web collection, laboratory data sharing, and screening of public datasets. After processing through standardized procedures such as Tibetan-Chinese machine translation, text normalization, speech synthesis, audio preprocessing, and six-tuple modality alignment, the final dataset contains six-tuple structured data including Tibetan spoken speech and text, Tibetan written speech and text, and Chinese text and speech. The TCST-6D-AD dataset includes one wav folder, one text folder, and Hexad-Metadata.json. The wav folder has three subfolders: tw_speech (Tibetan written speech), ts_speech (Tibetan spoken speech), m_speech (Chinese speech); the text folder contains three aligned text files: tw_text.tsv (Tibetan written text), ts_text.tsv (Tibetan spoken text), m_text.tsv (Chinese text).The overall scale of the dataset is as follows: it contains 10,068 six-tuple samples, with a total data size of 6.05 GB; among them, the text folder is 7.03 MB, the metadata file is 11 MB, and the audio files are 6.03 GB. The total audio durations are: Tibetan spoken speech 1,324.95 minutes, Tibetan written speech 1,165.44 minutes, and Chinese speech 883.84 minutes. We constructed the TCST-6D-AD end-to-end Tibetan-Chinese speech translation dataset. The source data was obtained through three channels: web collection, laboratory data sharing, and screening of public datasets. After processing through standardized procedures such as Tibetan-Chinese machine translation, text normalization, speech synthesis, audio preprocessing, and six-tuple modality alignment, the final dataset contains six-tuple structured data including Tibetan spoken speech and text, Tibetan written speech and text, and Chinese text and speech. The TCST-6D-AD dataset includes one wav folder, one text folder, and Hexad-Metadata.json. The wav folder has three subfolders: tw_speech (Tibetan written speech), ts_speech (Tibetan spoken speech), m_speech (Chinese speech); the text folder contains three aligned text files: tw_text.tsv (Tibetan written text), ts_text.tsv (Tibetan spoken text), m_text.tsv (Chinese text).The overall scale of the dataset is as follows: it contains 10,068 six-tuple samples, with a total data size of 6.05 GB; among them, the text folder is 7.03 MB, the metadata file is 11 MB, and the audio files are 6.03 GB. The total audio durations are: Tibetan spoken speech 1,324.95 minutes, Tibetan written speech 1,165.44 minutes, and Chinese speech 883.84 minutes.

Visit

doi.orgwww.scidb.cn

Tasks

speech translationspeech processingmachine translation

Tags

Computer science and technologyLow-resource languagesAmdo dialectTibetan-Chinese speech translationend-to-end speech translationsix-tuple dataset

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Parallel Corpus Sentence Alignment Scoring Dataset for Low-Resource Languages 低资源语言平行语料句对齐评分数据集Guidelines for Abstracts to be presented at the China Scientific Data 面向低资源语言教育与智能分析的中小学藏文作文数据集The dataset for "A Multilingual Short Text Classification Method Based on In-Context Learning" “基于上下文学习的多语言短文本情感分类方法”的数据集

Parallel Corpus Sentence Alignment Scoring Dataset for Low-Resource Languages 低资源语言平行语料句对齐评分数据集

Based on the neural network-based unsupervised sentence embedding method NeuroAlign, parallel senten

Guidelines for Abstracts to be presented at the China Scientific Data 面向低资源语言教育与智能分析的中小学藏文作文数据集

Spanning the period from 2011 to 2026, this dataset comprises a total of 1,187 essays. The corpus in

The dataset for "A Multilingual Short Text Classification Method Based on In-Context Learning" “基于上下文学习的多语言短文本情感分类方法”的数据集

In this paper, AfriSenti-SemEval is adopted as the experimental dataset. AfriSenti-SemEval is a data