Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Automatic transformation of Kazakh text into sign language glosses using multilingual transformer-based models

Domain:

natural language processing

Record type:

paper
Creator:
NurAigNazNaz
Publisher:
Fro
Host:
This study investigates the automatic transformation of Kazakh text into sign language glosses (Text-to-Gloss) using multilingual transformer-based models with emphasis on preserving morphological structure in a low-resource agglutinative language framework. Given the scarcity of high-quality intermediate representations for Kazakh Sign Language, a methodology for corpus formation was developed, resulting in a specialized dataset of 11 190 unique text−gloss pairs sourced from educational materials. To assess the effectiveness of transfer learning under low-resource constraints, we fine-tuned mT5-small and mT5-base models on this dataset. Experimental results indicate that the mT5-base architecture consistently outperforms the smaller variant across all metrics, achieving a 2.61-point improvement in BLEU (87.51 → 90.12) and a 2.47-point increase in ROUGE-1 F1 (89.61% → 92.08%), alongside gains in exact match (74.96% → 80.68%) and chrF (91.6 → 94.1). Statistical significance testing ( p < 0.001) confirms the stability of this improvement. Error analysis reveals that larger model capacity particularly improves morphological handling and token coverage, which is critical for agglutinative languages, while maintaining semantic accuracy. These findings demonstrate that pretrained multilingual models effectively address TTG transformation for low-resource agglutinative languages through transfer learning, with practical applications for Kazakh Sign Language assistive systems and education.

Visit

doi.org

Tasks

sign-language to textcomputer vision

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language ModelsERROR_500@DravidianLangTech2026: Automatic Prompt Style Classification in Telugu Using Transformer-Based Language ModelsPre-Trained Transformer-Based Models for Text Classification Using Low-Resourced Ewe LanguageMultilingual Polarization Detection Using Transformer-Based Models with Class Weighting and Threshold TuningEnhancing Automatic Speech Recognition for Moroccan Darija Using Transformer ModelsUsing Songs to Improve Kazakh Automatic Speech Recognition

PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models

Disinformation spreads rapidly across linguistic boundaries, yet most AI models are still benchmarke

ERROR_500@DravidianLangTech2026: Automatic Prompt Style Classification in Telugu Using Transformer-Based Language Models

Recovering writing style prompts in low resource languages has been daunting due to diverse morpholo

Pre-Trained Transformer-Based Models for Text Classification Using Low-Resourced Ewe Language

Despite a few attempts to automatically crawl Ewe text from online news portals and magazines, the A

Multilingual Polarization Detection Using Transformer-Based Models with Class Weighting and Threshold Tuning

This paper describes our submission to SemEval-2026 Task 9 on detecting multilingual, multicultural,

Enhancing Automatic Speech Recognition for Moroccan Darija Using Transformer Models

Using Songs to Improve Kazakh Automatic Speech Recognition

Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the