Logo Lanfrica

Language Identification in Low-Resource Multilingual Document Images using Deep Learning Techniques

Domain:

natural language processing

Record type:

paperdataset
Creator:
LomShuBir
Publisher:
MDP
Host:
With the increasing digitization of printed materials, it has become common to encounter documents in multiple languages in our daily lives. This has led to significant interest from researchers in the field of document image analysis, which includes applications like OCR, document image retrieval, and machine translation. However, these applications typically assume that documents are written in a single language, which is not always the case. Analyzing multilingual documents can be challenging, especially when the languages are switched within the document and the scripts are the same. In particular, this issue for low-resource languages has not yet been addressed. Hence, this paper focuses on exploring the use of deep learning techniques to develop a language identification model for multilingual documents written in Ethiopic script. To train and test our model, we extracted a total of 24,444 and 32,000 text-line pictures from a set of 814 and 1066 documents written in Ethiopian script and obtained from a variety of sources, respectively. We also carried out extensive experiments using both pre-trained CNN models and models trained from scratch with different parameter tuning on printed and synthetic document image datasets. The results revealed that CNN outperformed pre-trained deep learning models, achieving a promising accuracy of 97.27% on a test set of 9,600 images. These findings show the potential for developing effective language identification solutions for document images written in Ethiopic script.