Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Hybrid Modeling for Effective Text Spotting in Gujarati Language

Domain:

natural language processing

Record type:

paper
Creator:
DavDeg
Publisher:
Zenodo
Host:avatar
This paper presents a novel hybrid modeling approach for effective text spotting specifically tailored to the Gujarati language, achieving a high accuracy of 91.8% while maintaining efficient training time of only 18 minutes. The proposed hybrid model synergistically combines convolutional neural networks (CNN) for feature extraction and transformer-based architectures for contextual understanding, optimizing both recognition accuracy and computational efficiency. Gujarati, with its complex script and unique character shapes, presents challenges such as cursive and ligature forms, which the hybrid framework effectively addresses by leveraging spatial and sequential information jointly. Unlike traditional single-method models that either focus solely on spatial features or sequential patterns, our hybrid approach integrates both aspects to improve robustness against noise, background clutter, and varying text orientations in natural scene images. Experimental results on a custom-compiled dataset of Gujarati text in diverse scenes demonstrate superior performance compared to baseline models, with a notable reduction in false positives and recognition errors. The short training time also makes this method viable for real-world applications requiring quick model updates or deployment on resource-constrained devices. This work contributes a valuable advancement in Indic script OCR research and opens pathways for extending hybrid frameworks to other low-resource languages with complex scripts. The system's effectiveness in handling multi-style text and mixed backgrounds suggests promising potential for integration into mobile-based text reading applications and automated document processing for Gujarati text.

Visit

doi.org

Tasks

computer visionoptical character recognition

Tags

Text spotting; deep learning; scene text detection; optical character recognition (OCR); end-to-end frameworks; machine learning; computer vision

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Using ASR-Generated Text for Spoken Language ModelingHybrid Approach for Automatic Text Summarization for Low-resourced Amharic LanguageModeling for effective collaboration in telemedicineUsing web text to improve keyword spotting in speechEffective retrieval techniques for Arabic textEvaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

Using ASR-Generated Text for Spoken Language Modeling

International audience This papers aims at improving spoken language modeling (LM) us

Hybrid Approach for Automatic Text Summarization for Low-resourced Amharic Language

Automatic text summarization creates a concise version of the given document while retaining the ori

Modeling for effective collaboration in telemedicine

International audience Telemedicine is a remote medical practice, which utilizes adva

Using web text to improve keyword spotting in speech

For low resource languages, collecting sufficient training data to build acoustic and language model

Effective retrieval techniques for Arabic text

Arabic is a major international language, spoken in more than 23 countries, and the lingua franca of

Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware,