Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Low-Resource Speech-Driven NLP Pipeline for Sinhala Dyslexia Assistance

Domain:

natural language processinghealthcare

Record type:

paper
Creator:
PerSum
Host:avatar
Dyslexia in adults remains an under-researched and under-served area, particularly in non-English-speaking contexts, despite its significant impact on personal and professional lives. This work addresses that gap by focusing on Sinhala, a low-resource language with limited tools for linguistic accessibility. We present an assistive system explicitly designed for Sinhala-speaking adults with dyslexia. The system integrates Whisper for speech-to-text conversion, SinBERT, an open-sourced fine-tuned BERT model trained for Sinhala to identify common dyslexic errors, and a combined mT5 and Mistral-based model to generate corrected text. Finally, the output is converted back to speech using gTTS, creating a complete multimodal feedback loop. Despite the challenges posed by limited Sinhala-language datasets, the system achieves 0.66 transcription accuracy and 0.7 correction accuracy with 0.65 overall system accuracy. These results demonstrate both the feasibility and effectiveness of the approach. Ultimately, this work highlights the importance of inclusive Natural Language Processing (NLP) technologies in underrepresented languages and showcases a practical 11 pages, 4 figures, 3 tables

Visit

arxiv.org

Tasks

automatic speech recognitiongrammar error correctionspeech processingtext to speech

Tags

Computation and LanguageSoftware Engineering

Similar

A 10-Million-Row Sinhala Narrative Corpus for Low-Resource NLP: Dataset Construction, Statistical Characterisation, and Ethical Considerationsaizazayubi/Low-Resource-Speech-Dataset-Builder-Full-Pipeline-Decolonizing NLP for “Low-resource Languages”Design and Evaluation of a Scalable Data Pipeline for AI-Driven Air Quality Monitoring in Low-Resource SettingsA Sheng Phishing Corpus for Low-Resource Cybersecurity NLPBaaqar-007/dyslexia-accessibility-nlp

A 10-Million-Row Sinhala Narrative Corpus for Low-Resource NLP: Dataset Construction, Statistical Characterisation, and Ethical Considerations

sinhala_stories is a crowdsourced corpus of Sinhala-language narrative text comprising 10,949,004 ro

aizazayubi/Low-Resource-Speech-Dataset-Builder-Full-Pipeline-

# **Low Resource Speech Dataset Builder (Wav2Vec Friendly)** Tools for downloading speech from YouT

Decolonizing NLP for “Low-resource Languages”

Today African languages are spoken by more than a billion people, yet in the world of machine transl

Design and Evaluation of a Scalable Data Pipeline for AI-Driven Air Quality Monitoring in Low-Resource Settings

The increasing adoption of low-cost environmental sensors and AI-enabled applications has accelerate

A Sheng Phishing Corpus for Low-Resource Cybersecurity NLP

 

This dataset, the Sheng-English

Baaqar-007/dyslexia-accessibility-nlp

A multi-model ML pipeline that detects dyslexia indicators from handwriting images by combining a l