Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Handwritten Word Recognition for Low-Resource Languages: A CRNN-CTC Framework for Kirundi

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Niy
Publisher:
RSI
Host:
Handwritten Text Recognition (HTR) has experienced remarkable progress with the development of deep learning techniques. However, most existing studies focus on high-resource languages for which large annotated datasets are readily available. In contrast, low-resource languages remain largely underrepresented in handwriting recognition research due to the scarcity of handwritten corpora, linguistic resources, and benchmark datasets. This paper presents a handwritten word recognition framework for Kirundi, a low-resource Bantu language spoken primarily in Burundi. The proposed system employs a Convolutional Recurrent Neural Network (CRNN) combined with Connectionist Temporal Classification (CTC) for end-to-end sequence recognition without explicit character segmentation. To address severe data scarcity, a small handwritten Kirundi dataset consisting of manually collected word samples was constructed and annotated. Data augmentation techniques, including rotation, translation, Gaussian noise, Gaussian blur, and elastic distortion, were applied to increase sample diversity. In addition, synthetic handwritten-style data were generated to further expand the training set. Three experimental configurations were investigated: real handwritten data only, real data with augmentation, and real data combined with augmentation and synthetic handwritten-style images. Experimental results demonstrate that synthetic data generation improved recognition performance and reduced Character Error Rate (CER) from 0.8354 to 0.7560, corresponding to an approximate relative improvement of 9.5%. Although exact word-level recognition remained difficult because of the extremely limited dataset size, the proposed framework successfully learned meaningful sequential patterns and produced increasingly structured Kirundi-like predictions. The study establishes an initial benchmark for Kirundi handwritten word recognition and highlights the potential of synthetic data generation for low-resource handwriting recognition tasks.

Visit

doi.org

Tasks

computer visionoptical character recognition

Languages

Rundi

Similar

CRNN-CTC Deep Learning Algorithm for Handwritten Amharic Document Recognitionjoearid/Amharic-OCR-CRNN-CTC-for-Printed-Ge-ez-Script-RecognitionFuzzy Crab Optimization Based Framework for Handwritten Character Recognition in Low Resource ScriptsCRNN-based Word Recognition Model for Reading Algerian Drug LabelsRobust speech recognition for low-resource languagesTest-Time Adaptation for Low-Resource Handwritten Character Recognition under Distribution Shift

CRNN-CTC Deep Learning Algorithm for Handwritten Amharic Document Recognition

Preserving handwritten Amharic documents for long periods of time requires converting them into a co

joearid/Amharic-OCR-CRNN-CTC-for-Printed-Ge-ez-Script-Recognition

Amharic OCR: CRNN-CTC pipeline for printed Ge'ez script. 96.3% accuracy, 0.64% CER. CNN+BiLSTM+CTC w

Fuzzy Crab Optimization Based Framework for Handwritten Character Recognition in Low Resource Scripts

Handwritten character recognition in low-resource scripts is a challenging yet critical task for pre

CRNN-based Word Recognition Model for Reading Algerian Drug Labels

Robust speech recognition for low-resource languages

Process of human-machine interaction is an integral part of everyday human life in a modern world. T

Test-Time Adaptation for Low-Resource Handwritten Character Recognition under Distribution Shift

This study presents the first systematic evaluation of test-time adaptation for low-resource handwri