Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

An approach to analysis of arabic text documents into text lines, words, and characters

Domain:

natural language processing

Record type:

paper
Creator:
HakAhmRamSho
Publisher:
Institute of Advanced Engineering and Science
Host:
Text line extraction from a text document image and segmenting it into isolate words and segmenting these words into individual characters are considered as one of the most critical processes in OCR systems development and turning the document into a searchable electronic representation, this paper presents a new approach to analyze the Arabic text documents, the proposed approach contains four steps, preprocessing, text line segmentation, word segmentation, character segmentation. The horizontal projection method are used to detect and extract the text line from preprocessed text documents image, in word segmentation step The space threshold are computed to determine the spaces among connected components in text line as within-word space or between-words space for segmenting the text line into isolate words, finally thinning method applied to find the skeleton of segmented word and analyses geometric characteristics of the characters to detect ligatures and characters. The proposed approach was tested and evaluated on a set of 115 text images, this set contains images from the KFUPM Handwritten Arabic TexT (KHATT) database and some images produced by the authors. The experiment results are extremely encouraging, with a success rate of 98.6% for lines segmentation, 96% for words segmentation, and 87.1% for characters segmentation.

Visit

doi.org

Tasks

computer visionoptical character recognition

Licenses

http://creativecommons.org/licenses/by-nc/4.0

Similar

Design and Development of an Arabic Text-To-Speech SynthesizerGraph-Based Text Modeling: Considering Mathematical Semantic Linking to Improve the Indexation of Arabic DocumentsAn Automatic Stop Words Removal in Maghrebi Arabic Dialect Text Classification Using Part of Speech TaggingAutomatic detection of English words in Benglish text: A statistical approachNatiQ: An End-to-end Text-to-Speech System for ArabicAn NLP based text-to-speech synthesizer for Moroccan Arabic

Design and Development of an Arabic Text-To-Speech Synthesizer

Computers that can interact with humans via speech had become a dream for scientists since the early

Graph-Based Text Modeling: Considering Mathematical Semantic Linking to Improve the Indexation of Arabic Documents

International audience

An Automatic Stop Words Removal in Maghrebi Arabic Dialect Text Classification Using Part of Speech Tagging

Automatic detection of English words in Benglish text: A statistical approach

NatiQ: An End-to-end Text-to-Speech System for Arabic

NatiQ is end-to-end text-to-speech system for Arabic. Our speech synthesizer uses an encoder-decoder

An NLP based text-to-speech synthesizer for Moroccan Arabic