Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Robustness of Synthetic vs. Human-Annotated Grammatical Error Detection Models in Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, these annotations are unavailable in many low-resource languages. In this paper, we investigate GED in this context. Leveraging the zero-shot cross-lingual transfer capabilities of multilingual pre-trained language models, we train a model using data from a diverse set of languages to generate synthetic errors in other languages. These synthetic error corpora are then used to train a GED model. Specifically we propose a two-stage fine-tuning pipeline where the GED model is first fine-tuned on mult Research goal: How does the robustness of grammatical error detection models trained on zero-shot synthetic data vary against adversarial noise compared to models trained on human-annotated corpora in low-resource FLORES-200 languages? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.0/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 9.0/10.

Visit

doi.orgzenodo.org

Tasks

grammar error correction

Tags

robustnessgrammaticalerrordetectionmodelstrainedzero-shotsynthetic

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Grammatical Error Detection Performance in Low-Resource Languages: Zero-Shot Synthetic vs. Human-Annotated BaselinesCorrelation between Source Language Diversity and Synthetic Data Robustness in Low-Resource Grammatical Error DetectionDiversity in Zero-Shot Synthetic Data for Low-Resource Grammatical Error DetectionSynthetic Data Diversity and Robustness in Teacher-Student NER Models for Low-Resource Languages

Grammatical Error Detection Performance in Low-Resource Languages: Zero-Shot Synthetic vs. Human-Annotated Baselines

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Correlation between Source Language Diversity and Synthetic Data Robustness in Low-Resource Grammatical Error Detection

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Diversity in Zero-Shot Synthetic Data for Low-Resource Grammatical Error Detection

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Synthetic Data Diversity and Robustness in Teacher-Student NER Models for Low-Resource Languages

Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for language