Logo Lanfrica

Tibetan-STR Tibetan-STR

Domain:

natural language processing

Record type:

dataset
Creator:
NimGaoSonFei
Publisher:
Sci
Host:avatar
Tibetan, as a typical low-resource language, has long lacked publicly available datasets in the field of scene text recognition (STR). This absence not only restricts the iteration of traditional recognition algorithms but also makes it difficult to evaluate the generalization capabilities and fine-tuning effects of large vision-language models (VLMs) on Tibetan. To this end, this paper constructs Tibetan-STR, a natural scene image dataset of Tibetan Uchen script. The dataset contains 4,322 real-scene images covering diverse scenarios such as commercial signboards, traffic signs, and public notices, providing precise text annotations. The images deeply couple the visual degradation caused by the extreme physical environment of the plateau with the unique 2D non-linear vertical stacking orthographic features of Tibetan script, effectively capturing the real-world challenges of Tibetan scene text recognition.The dataset is stratified and randomly sampled according to Tibetan syllable length and stacked character density, and split into a training set (3,458 images), a validation set (432 images), and a test set (432 images) at an 8:1:1 ratio. The corresponding splitting scripts and index files are released together with the data. To validate the dataset and provide baselines, this paper conducts a systematic evaluation using representative models across different technical paradigms. These include the traditional 1D sequence model CRNN, the 2D vision Transformer model SVTRv2, and the large vision-language model PaddleOCR-VL. Additionally, the syllable error rate (SER) is introduced as an evaluation metric tailored to Tibetan orthographic characteristics. Baseline experimental results show that the best-performing model, SVTRv2, achieves a line-level accuracy of only 61.34%, revealing that significant performance challenges remain in complex Tibetan scenes. The Tibetan-STR dataset provides a standardized and reproducible public benchmark for low-resource language scene text recognition, offering important data support for advancing multilingual OCR research. Tibetan, as a typical low-resource language, has long lacked publicly available datasets in the field of scene text recognition (STR). This absence not only restricts the iteration of traditional recognition algorithms but also makes it difficult to evaluate the generalization capabilities and fine-tuning effects of large vision-language models (VLMs) on Tibetan. To this end, this paper constructs Tibetan-STR, a natural scene image dataset of Tibetan Uchen script. The dataset contains 4,322 real-scene images covering diverse scenarios such as commercial signboards, traffic signs, and public notices, providing precise text annotations. The images deeply couple the visual degradation caused by the extreme physical environment of the plateau with the unique 2D non-linear vertical stacking orthographic features of Tibetan script, effectively capturing the real-world challenges of Tibetan scene text recognition.The dataset is stratified and randomly sampled according to Tibetan syllable length and stacked character density, and split into a training set (3,458 images), a validation set (432 images), and a test set (432 images) at an 8:1:1 ratio. The corresponding splitting scripts and index files are released together with the data. To validate the dataset and provide baselines, this paper conducts a systematic evaluation using representative models across different technical paradigms. These include the traditional 1D sequence model CRNN, the 2D vision Transformer model SVTRv2, and the large vision-language model PaddleOCR-VL. Additionally, the syllable error rate (SER) is introduced as an evaluation metric tailored to Tibetan orthographic characteristics. Baseline experimental results show that the best-performing model, SVTRv2, achieves a line-level accuracy of only 61.34%, revealing that significant performance challenges remain in complex Tibetan scenes. The Tibetan-STR dataset provides a standardized and reproducible public benchmark for low-resource language scene text recognition, offering important data support for advancing multilingual OCR research.