Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Towards the extraction of robust sign embeddings for low resource sign language recognition

Domain:

natural language processing

Record type:

papermodel
Creator:
De RusHolVen
Host:avatar
Isolated Sign Language Recognition (SLR) has mostly been applied on datasets containing signs executed slowly and clearly by a limited group of signers. In real-world scenarios, however, we are met with challenging visual conditions, coarticulated signing, small datasets, and the need for signer independent models. To tackle this difficult problem, we require a robust feature extractor to process the sign language videos. One could expect human pose estimators to be ideal candidates. However, due to a domain mismatch with their training sets and challenging poses in sign language, they lack robustness on sign language data and image-based models often still outperform keypoint-based models. Furthermore, whereas the common practice of transfer learning with image-based models yields even higher accuracy, keypoint-based models are typically trained from scratch on every SLR dataset. These factors limit their usefulness for SLR. From the existing literature, it is also not clear which, if any, pose estimator performs best for SLR. We compare the three most popular pose estimators for SLR: OpenPose, MMPose and MediaPipe. We show that through keypoint normalization, missing keypoint imputation, and learning a pose embedding, we can obtain significantly better results and enable transfer learning. We show that keypoint-based embeddings contain cross-lingual features: they can transfer between sign languages and achieve competitive performance even when fine-tuning only the classifier layer of an SLR model on a target sign language. We furthermore achieve better performance using fine-tuned transferred embeddings than models trained only on the target sign language. The embeddings can also be learned in a multilingual fashion. The application of these embeddings could prove particularly useful for low resource sign languages in the future.

Visit

arxiv.org

Tasks

computer visionembeddingssign-language to text

Tags

Computer Vision and Pattern RecognitionComputation and Language

Similar

Prompting with Sign Parameters for Low-resource Sign Language Instruction GenerationTowards Kenyan Sign Language Hand Gesture Recognition DatasetA Generic Approach towards Amharic Sign Language RecognitionSpatial-Temporal Feature Extraction for Tanzanian Sign Language Recognition in Medical DiagnosticsAmharic Static Alphabet Sign Language Recognition Based on Hybrid Feature Extraction ApproachEvaluating the Impact of Deep Learning Model Architecture on Sign Language Recognition Accuracy in Low-Resource Context

Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation

Sign Language (SL) enables two-way communication for the deaf and hard-of-hearing community, yet man

Towards Kenyan Sign Language Hand Gesture Recognition Dataset

Datasets for hand gesture recognition are now an important aspect of machine learning. Many datasets

A Generic Approach towards Amharic Sign Language Recognition

In the day-to-day life of communities, good communication channels are crucial for mutual understand

Spatial-Temporal Feature Extraction for Tanzanian Sign Language Recognition in Medical Diagnostics

Amharic Static Alphabet Sign Language Recognition Based on Hybrid Feature Extraction Approach

Sign language is a communication mechanism for hearing-impaired people. Unless hearing individuals l

Evaluating the Impact of Deep Learning Model Architecture on Sign Language Recognition Accuracy in Low-Resource Context

Deep learning models are well-known for their reliance on large training datasets to achieve optimal