Logo Lanfrica

Developing Automatic Character Recognition System for Guragigna Printed Real-Life Document

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
Tse
Éditeur:
Sul
Éditeur:
Nat
Hôte:avatar
Optical Character Recognition is an area of research and development where a system is made to
recognize characters from printed documents. Cultural considerations and enormous flood of
printed documents motivated the development of Optical Character Recognition across the
world. However; Guragigna character recognition still needs to enhance its performance, didn"t
cover all sizzling issues and there is a lack of well-designed preprocessing method. In this
proposed system the recognition supports multi-font and style of Guragigna characters which
gives an enhancement for the performance of the system; therefore, it is an advanced Optical
Character Recognition system. This thesis work presents different machine learning approach for
preprocessing, segmentation and recognition of Guragigna script document. The study also uses
python programming language and we utilize packages or libraries such as OpenCV, scikitimage, scikit-learn and keras. For the recognition to be successful, the design of character
recognition is implemented by using convolution neural network.
This experiment trained on 12,018 total samples, validated with 3653 samples. The final result
implies that the learning in the network improved and can generalize for the new test data with
higher recognition rate. 98.89% validation accuracy recorded for the augmented data and 99.86%
is recorded for live-split data. After all, we can record 99.75% total training accuracy. Our
contribution and advantage of work include; the learning process with an approach for
Guragigna character recognition system, semi-automatic techniques developed to solve data
labeling, and the scheme could easily be applied to other similar with little effort. In addition,
the dataset preparation, preprocessing, segmentation, annotation and classification for
recognition techniques are demonstrated through the experimental result with standard
evaluation. Extension works are recommended that need further consideration in preprocessing,
segmentation, recognition and enhancement of document images. In addition, the use of machine
learning techniques to perform an automatic segmentation and labeling with the usage of some
structural information to raise the performance for better performance of the study.