Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

New Results for the Text Recognition of Arabic Maghrib{ī} Manuscripts -- Managing an Under-resourced Script

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
NoëSalVid
Hôte:avatar
HTR models development has become a conventional step for digital humanities projects. The performance of these models, often quite high, relies on manual transcription and numerous handwritten documents. Although the method has proven successful for Latin scripts, a similar amount of data is not yet achievable for scripts considered poorly-endowed, like Arabic scripts. In that respect, we are introducing and assessing a new modus operandi for HTR models development and fine-tuning dedicated to the Arabic Maghrib{ī} scripts. The comparison between several state-of-the-art HTR demonstrates the relevance of a word-based neural approach specialized for Arabic, capable to achieve an error rate below 5% with only 10 pages manually transcribed. These results open new perspectives for Arabic scripts processing and more generally for poorly-endowed languages processing. This research is part of the development of RASAM dataset in partnership with the GIS MOMM and the BULAC.

Visit

arxiv.org

Tasks

computer visionoptical character recognition

Tags

Computer Vision and Pattern Recognition

Similaires

Automatic speech recognition for an under-resourced language - amharicScript Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual CommunitiesShort Text Language Identification for Under Resourced LanguagesInvestigating Data Sharing in Speech Recognition for an Under-Resourced Language: The Case of Algerian DialectAutomatic speech recognition for under-resourced languages: A surveySub-word Based End-to-End Speech Recognition for an Under-Resourced Language: Amharic

Automatic speech recognition for an under-resourced language - amharic

Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities

The wide accessibility of social media has provided linguistically under-represented communities wit

Short Text Language Identification for Under Resourced Languages

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text languag

Investigating Data Sharing in Speech Recognition for an Under-Resourced Language: The Case of Algerian Dialect

The Arabic language has many varieties, including its standard form, Modern Standard Arabic (MSA), a

Automatic speech recognition for under-resourced languages: A survey

(Impact-F 1.28 estim. in 2012) International audience no abstract

Sub-word Based End-to-End Speech Recognition for an Under-Resourced Language: Amharic

In this work, we focused on end-to-end speech recognition for less-resourced language, Amharic. The result can be integrated with other tasks such as spoken content retrieval. We explored three models, which consist of Convolutional Neural Networks, Recurrent Neura