Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Tibetan-STR Tibetan-STR

Domaine:

natural language processing

Type de record:

dataset
Créateur:
NimGaoSonFei
Éditeur:
Sci
Hôte:avatar
Tibetan, as a typical low-resource language, has long lacked publicly available datasets in the field of scene text recognition (STR). This absence not only restricts the iteration of traditional recognition algorithms but also makes it difficult to evaluate the generalization capabilities and fine-tuning effects of large vision-language models (VLMs) on Tibetan. To this end, this paper constructs Tibetan-STR, a natural scene image dataset of Tibetan Uchen script. The dataset contains 4,322 real-scene images covering diverse scenarios such as commercial signboards, traffic signs, and public notices, providing precise text annotations. The images deeply couple the visual degradation caused by the extreme physical environment of the plateau with the unique 2D non-linear vertical stacking orthographic features of Tibetan script, effectively capturing the real-world challenges of Tibetan scene text recognition.The dataset is stratified and randomly sampled according to Tibetan syllable length and stacked character density, and split into a training set (3,458 images), a validation set (432 images), and a test set (432 images) at an 8:1:1 ratio. The corresponding splitting scripts and index files are released together with the data. To validate the dataset and provide baselines, this paper conducts a systematic evaluation using representative models across different technical paradigms. These include the traditional 1D sequence model CRNN, the 2D vision Transformer model SVTRv2, and the large vision-language model PaddleOCR-VL. Additionally, the syllable error rate (SER) is introduced as an evaluation metric tailored to Tibetan orthographic characteristics. Baseline experimental results show that the best-performing model, SVTRv2, achieves a line-level accuracy of only 61.34%, revealing that significant performance challenges remain in complex Tibetan scenes. The Tibetan-STR dataset provides a standardized and reproducible public benchmark for low-resource language scene text recognition, offering important data support for advancing multilingual OCR research. Tibetan, as a typical low-resource language, has long lacked publicly available datasets in the field of scene text recognition (STR). This absence not only restricts the iteration of traditional recognition algorithms but also makes it difficult to evaluate the generalization capabilities and fine-tuning effects of large vision-language models (VLMs) on Tibetan. To this end, this paper constructs Tibetan-STR, a natural scene image dataset of Tibetan Uchen script. The dataset contains 4,322 real-scene images covering diverse scenarios such as commercial signboards, traffic signs, and public notices, providing precise text annotations. The images deeply couple the visual degradation caused by the extreme physical environment of the plateau with the unique 2D non-linear vertical stacking orthographic features of Tibetan script, effectively capturing the real-world challenges of Tibetan scene text recognition.The dataset is stratified and randomly sampled according to Tibetan syllable length and stacked character density, and split into a training set (3,458 images), a validation set (432 images), and a test set (432 images) at an 8:1:1 ratio. The corresponding splitting scripts and index files are released together with the data. To validate the dataset and provide baselines, this paper conducts a systematic evaluation using representative models across different technical paradigms. These include the traditional 1D sequence model CRNN, the 2D vision Transformer model SVTRv2, and the large vision-language model PaddleOCR-VL. Additionally, the syllable error rate (SER) is introduced as an evaluation metric tailored to Tibetan orthographic characteristics. Baseline experimental results show that the best-performing model, SVTRv2, achieves a line-level accuracy of only 61.34%, revealing that significant performance challenges remain in complex Tibetan scenes. The Tibetan-STR dataset provides a standardized and reproducible public benchmark for low-resource language scene text recognition, offering important data support for advancing multilingual OCR research.

Visit

doi.orgwww.scidb.cn

Tags

Information science and systems scienceBasic discipline of engineering and technologyInformation and systems science related engineering and technologyNatural science related engineering and technologyComputer science and technologyTibetan-STR DatasetNatural Scene ImagesTibetan Uchen ScriptSyllable Error Rate (SER)Low-Resource Language OCR

Licenses

Creative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcode

Similaires

Developing the Old Tibetan TreebankTibetan-PASEM: Phonology-Aware Speech Evidence Matching for Low-Resource Tibetan Written-Query Keyword SpottingFadilaW/Swahili-STR-DatasetTiBERT: Tibetan Pre-trained Language Model17 STR data (AmpF/STR Identifiler and Powerplex 16 System) from Cabinda (Angola)Empis Shamshev, 2020, s. str.

Developing the Old Tibetan Treebank

This paper presents a full procedure for the development of a segmented, POS-tagged and chunk-parsed

Tibetan-PASEM: Phonology-Aware Speech Evidence Matching for Low-Resource Tibetan Written-Query Keyword Spotting

Low-resource written-query keyword spotting detects a text-specified target in speech without spoken

FadilaW/Swahili-STR-Dataset

Swahili Language Scene Text Detection and Recognition Dataset # Swahili-STR-Dataset Accepted by I

TiBERT: Tibetan Pre-trained Language Model

The pre-trained language model is trained on large-scale unlabeled text and can achieve state-of-the

17 STR data (AmpF/STR Identifiler and Powerplex 16 System) from Cabinda (Angola)

Empis Shamshev, 2020, s. str.

Key to species of Empis s. str. of Egypt, Israel and Syria 1 Male ..................................