Logo Lanfrica

Contribution to Arabic Character Recognition Based on Directional Approaches CONTRIBUTION A LA RECONNAISSANCE DES CARACTERES ARABES PAR APPROCHES DIRECTIONNELLES

Domaine:

natural language processing

Type de record:

paper
Créateur:
Kad
Éditeur:
FacFacM’b
Éditeur:
CCSD
Hôte:avatar
Digital recognition of Arabic characters written in real time or contained in a captured image of handwritten or printed text is a difficult problem that still exists. Many methods and approaches are proposed to solve this problem, but the results so far remain unsatisfactory. Many of these methods were originally designed to recognize Latin and Chinese characters. They do not suit the nature of Arabic writing which is usually cursive and curved or need to be modified to accommodate Arabic characters. In this thesis, we propose a new approach to identify Arabic characters. Considering the curved nature of Arabic writing, we see that the methods which provide information about the changes and directions of the character curve are the most appropriate. We have chosen two methods according to this vision : the Freeman coding, and the Hough transform. Then we improved each method and modified it to make it more suitable for Arabic characters.The main improvement we made was to introduce a new algorithm which can give a short and precise Freeman code for the Arabic character in all its possible forms. And compared to the traditional algorithm, there has been a noticeable improvement in the results of recognizing Arabic characters after using the innovative algorithm in producing the code along with other proposed methods to generalize and normalize the code. As for the method based on the Hough transform, we have proposed an improvement by using two thresholds to also take the low values in the accumulator matrix. This is far from the traditional method which focuses on peaks in the accumulator matrix and uses a single threshold. And after a preliminary classification of the Arabic character according to the number of loops, the nature, number and position of complementary parts (Points, Hamza, etc.), the results were remarkable. Especially the ability to recognize other different fonts which were not learned in the training phase and with a small number of learned fonts. La reconnaissance numérique des caractères arabes écrits en temps réel ou contenus dans une image capturée d’un texte manuscrit ou imprimé est un problème difficile qui existe toujours. De nombreuses méthodes et approches sont proposées pour résoudre ce problème, mais les résultats restent jusqu’à présent insatisfaisants. Beaucoup de ces méthodes ont été conçues à l’origine pour reconnaître les caractères latins et chinois. Elles ne conviennent pas à la nature de l’écriture arabe qui est généralement cursive et courbée ou doivent être modifiées pour s’adapter aux caractères arabes. Dans cette thèse, nous proposons une nouvelle approche pour identifier les caractères arabes. Compte tenu de la nature courbée de l’écriture arabe, nous voyons que les méthodes qui fournissent des informations sur les changements et les directions de la courbe du caractère sont les plus appropriées. Nous avons choisi deux méthodes selon cette vision : le codage de Freeman et la transformée de Hough. Ensuite, nous avons amélioré chaque méthode et l’avons modifiée pour la rendre plus adaptée aux caractères arabes. La plus grande amélioration que nous avons faite a été d’introduire un nouvel algorithme qui peut donner un code de Freeman court et précis au caractère arabe sous toutes ses formes possibles. Et par rapport à l’algorithme traditionnel, il y a eu une amélioration notoire dans les résultats de la reconnaissance des caractères arabes après avoir utilisé l’algorithme innovant dans la production du code avec d’autres méthodes proposées pour généraliser et normaliser le code. Quant à la méthode basée sur la transformée de Hough, nous avons proposé une amélioration en utilisant deux seuils pour prendre également les valeurs les plus basses dans la matrice d’accumulation. C’est loin de la méthode traditionnelle qui se concentre sur les pics dans la matrice d’accumulation et utilise un seuil unique. Et après une classification préliminaire du caractère arabe en fonction du nombre de boucles, de la nature, du nombre et de la position des parties complémentaires (Points, Hamza, etc.), les résultats étaient remarquables. En particulier la capacité de reconnaître d’autres fontes différentes qui n’ont pas été acquises dans l’étape d’apprentissage et avec un petit nombre de fontes apprises.

Languages

Similaires