Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Dynamic decoding and dual synthetic data for automatic correction of grammar in low-resource scenario

Domaine:

natural language processing

Type de record:

paper
Créateur:
AhmYinAimSir
Éditeur:
Pee
Hôte:
Grammar error correction systems are pivotal in the field of natural language processing (NLP), with a primary focus on identifying and correcting the grammatical integrity of written text. This is crucial for both language learning and formal communication. Recently, neural machine translation (NMT) has emerged as a promising approach in high demand. However, this approach faces significant challenges, particularly the scarcity of training data and the complexity of grammar error correction (GEC), especially for low-resource languages such as Indonesian. To address these challenges, we propose InSpelPoS, a confusion method that combines two synthetic data generation methods: the Inverted Spellchecker and Patterns+POS. Furthermore, we introduce an adapted seq2seq framework equipped with a dynamic decoding method and state-of-the-art Transformer-based neural language models to enhance the accuracy and efficiency of GEC. The dynamic decoding method is capable of navigating the complexities of GEC and correcting a wide range of errors, including contextual and grammatical errors. The proposed model leverages the contextual information of words and sentences to generate a corrected output. To assess the effectiveness of our proposed framework, we conducted experiments using synthetic data and compared its performance with existing GEC systems. The results demonstrate a significant improvement in the accuracy of Indonesian GEC compared to existing methods.

Visit

doi.org

Tasks

grammar error correction

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Semi-supervised learning and bidirectional decoding for effective grammar correction in low-resource scenariosSynthetic Data and Annotation Projection for Low-Resource NER PerformanceDiversity in Zero-Shot Synthetic Data for Low-Resource Grammatical Error DetectionSynthetic Data Diversity and Robustness in Teacher-Student NER Models for Low-Resource LanguagesSynthetic Data Diversity vs. Back-Translation for Multilingual NER in Low-Resource Languagestabularis-ai/Synthetic-Data-Generation-Pipeline-for-Low-Resource-Swahili-Sentiment-Analysis

Semi-supervised learning and bidirectional decoding for effective grammar correction in low-resource scenarios

The correction of grammatical errors in natural language processing is a crucial task as it aims to

Synthetic Data and Annotation Projection for Low-Resource NER Performance

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Diversity in Zero-Shot Synthetic Data for Low-Resource Grammatical Error Detection

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Synthetic Data Diversity and Robustness in Teacher-Student NER Models for Low-Resource Languages

Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for language

Synthetic Data Diversity vs. Back-Translation for Multilingual NER in Low-Resource Languages

Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for language

tabularis-ai/Synthetic-Data-Generation-Pipeline-for-Low-Resource-Swahili-Sentiment-Analysis

# Synthetic Data for Low-Resource Swahili Language Sentiment Analysis This repository contains the