Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Guidelines for Abstracts to be presented at the China Scientific Data 面向低资源语言教育与智能分析的中小学藏文作文数据集

Domain:

natural language processingeducation

Record type:

dataset
Creator:
Xid
Publisher:
Sci
Host:avatar
Spanning the period from 2011 to 2026, this dataset comprises a total of 1,187 essays. The corpus integrates content from both print publications and online resources, encompassing authoritative works—such as the *Selected Essays by Tibetan Primary School Students*—as well as texts sourced from educational websites. During the data acquisition process, Optical Character Recognition (OCR) technology was employed to digitize print texts, supplemented by manual proofreading to correct recognition errors and formatting issues. Online materials were acquired through a combined approach of targeted web crawling and manual curation, thereby ensuring the authenticity and diversity of the data. The dataset features six structured fields: essay ID, title, thematic category, educational stage, author, and body text. Regarding thematic annotation, the essays are categorized into six groups—including "Nature and Daily Life" and "Campus Life and Growth"—based on the distribution of humanistic themes found in the standardized Chinese language textbooks prescribed by the Ministry of Education.The dataset consists of one Excel (.xlsx) file, in which: (1) the file contains three worksheets corresponding to three educational stages: senior high school, junior high school, and primary school; (2) each worksheet includes five columns: Composition ID, Title, Topic Category, Educational Stage, and Full Text. Spanning the period from 2011 to 2026, this dataset comprises a total of 1,187 essays. The corpus integrates content from both print publications and online resources, encompassing authoritative works—such as the *Selected Essays by Tibetan Primary School Students*—as well as texts sourced from educational websites. During the data acquisition process, Optical Character Recognition (OCR) technology was employed to digitize print texts, supplemented by manual proofreading to correct recognition errors and formatting issues. Online materials were acquired through a combined approach of targeted web crawling and manual curation, thereby ensuring the authenticity and diversity of the data. The dataset features six structured fields: essay ID, title, thematic category, educational stage, author, and body text. Regarding thematic annotation, the essays are categorized into six groups—including "Nature and Daily Life" and "Campus Life and Growth"—based on the distribution of humanistic themes found in the standardized Chinese language textbooks prescribed by the Ministry of Education.The dataset consists of one Excel (.xlsx) file, in which: (1) the file contains three worksheets corresponding to three educational stages: senior high school, junior high school, and primary school; (2) each worksheet includes five columns: Composition ID, Title, Topic Category, Educational Stage, and Full Text.

Visit

doi.orgwww.scidb.cn

Tags

LinguisticsFOS: Languages and literaturePedagogyComputer science and technologyPrimary and Secondary School StudentsTibetan CompositionsDatasetLow-resource

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-sa/4.0/legalcode

Similar

Parallel Corpus Sentence Alignment Scoring Dataset for Low-Resource Languages 低资源语言平行语料句对齐评分数据集The dataset for "A Multilingual Short Text Classification Method Based on In-Context Learning" “基于上下文学习的多语言短文本情感分类方法”的数据集TCST-6D-AD: A Six-tuple Parallel Tibetan-Chinese Speech Translation Dataset for the Amdo Dialect TCST-6D-AD: 面向安多方言的藏汉六元组平行语音翻译数据集Replication Data for: 联合国维和部队派遣国构成与平民保护 ——基于非洲维和特派团的微观数据分析

Parallel Corpus Sentence Alignment Scoring Dataset for Low-Resource Languages 低资源语言平行语料句对齐评分数据集

Based on the neural network-based unsupervised sentence embedding method NeuroAlign, parallel senten

The dataset for "A Multilingual Short Text Classification Method Based on In-Context Learning" “基于上下文学习的多语言短文本情感分类方法”的数据集

In this paper, AfriSenti-SemEval is adopted as the experimental dataset. AfriSenti-SemEval is a data

TCST-6D-AD: A Six-tuple Parallel Tibetan-Chinese Speech Translation Dataset for the Amdo Dialect TCST-6D-AD: 面向安多方言的藏汉六元组平行语音翻译数据集

We constructed the TCST-6D-AD end-to-end Tibetan-Chinese speech translation dataset. The source data

Replication Data for: 联合国维和部队派遣国构成与平民保护 ——基于非洲维和特派团的微观数据分析

There is a growing body of scholarship on the effectiveness of UN Peacekeeping in protecting civilia