Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Hybrid Deep Learning Models for Arabic Sign Language Recognition in Healthcare Applications

Domain:

healthcarenatural language processing

Record type:

paperdataset
Creator:
ManHamLajAba
Editor:
InnSyn3iLMul
Publisher:
CCSDMDPI
Host:avatar
International audience Deaf and hearing-impaired individuals rely on sign language, a visual communication system using hand shapes, facial expressions, and body gestures. Sign languages vary by region. For example, Arabic Sign Language (ArSL) is notably different from American Sign Language (ASL). This project focuses on creating an Arabic Sign Language Recognition (Ar-SLR) System tailored for healthcare, aiming to bridge communication gaps resulting from a lack of sign-proficient professionals and limited region-specific technological solutions. Our research addresses limitations in sign language recognition systems by introducing a novel framework centered on ResNet50ViT, a hybrid architecture that synergistically combines ResNet50's robust local feature extraction with the global contextual modeling of Vision Transformers (ViT). We also explored a tailored Vision Transformer variant (SignViT) for Arabic Sign Language as a comparative model. Our main contribution is the ResNet50ViT model, which significantly outperforms existing approaches, specifically targeting the challenges of capturing sequential hand movements, which traditional CNN-based methods struggle with. We utilized an extensive dataset incorporating both static (36 signs) and dynamic (92 signs) medical signs. Through targeted preprocessing techniques and optimization strategies, we achieved significant performance improvements over conventional approaches. In our experiments, the proposed ResNet50-ViT achieved a remarkable 99.86% accuracy on the ArSL dataset, setting a new state-of-the-art, demonstrating the effectiveness of integrating ResNet50's hierarchical local feature extraction with Vision Transformer's global contextual modeling. For comparison, a fine-tuned Vision Transformer (SignViT) attained 98.03% accuracy, confirming the strength of transformer-based approaches but underscoring the clear performance gain enabled by our hybrid architecture. We expect that RAFID will help deaf patients communicate better with healthcare providers without needing human interpreters.

Visit

hal.science

Tasks

computer visionsign-language to text

Tags

ArSLrecognitiondeep learningfine-tuningCNNvision transformer[INFO]Computer Science [cs][INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]

Licenses

https://creativecommons.org/licenses/by/4.0/info:eu-repo/semantics/OpenAccess

Similar

A Hybrid Deep Learning Model for Arabic Text RecognitionEfficient Deep Learning Algorithm for Egyptian Sign Language RecognitionDeep Learning Shape Trajectories for Isolated Word Sign Language RecognitionMoroccan Sign Language Video Recognition with Deep LearningDeep Learning-Based Arabic Sign Language Digit Recognition Using ConvNeXt Architecture with Focal Loss OptimizationSign Language Recognition Using Deep Learning: Advancements and Challenges

A Hybrid Deep Learning Model for Arabic Text Recognition

Arabic text recognition is a challenging task because of the cursive nature of Arabic writing system

Efficient Deep Learning Algorithm for Egyptian Sign Language Recognition

Deep Learning Shape Trajectories for Isolated Word Sign Language Recognition

In this paper, we propose an efficient trajectories analysis solution for the recognition of Isolate

Moroccan Sign Language Video Recognition with Deep Learning

Deep Learning-Based Arabic Sign Language Digit Recognition Using ConvNeXt Architecture with Focal Loss Optimization

Sign Language Recognition Using Deep Learning: Advancements and Challenges

Abstract: Sign language recognition (SLR) has arisen as a major area of research in recent years, at