Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Performance comparison of synthetic error data via zero-shot cross-lingual transfer and human-annotated data on CoNLL-14 benchmark

Domain:

natural language processing

Record type:

paper
Creator:
Ass
Publisher:
Zenodo
Host:avatar
Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, these annotations are unavailable in many low-resource languages. In this paper, we investigate GED in this context. Leveraging the zero-shot cross-lingual transfer capabilities of multilingual pre-trained language models, we train a model using data from a diverse set of languages to generate synthetic errors in other languages. These synthetic error corpora are then used to train a GED model. Specifically we propose a two-stage fine-tuning pipeline where the GED model is first fine-tuned on mult Research goal: How does the performance of synthetic error data generated via zero-shot cross-lingual transfer compare to human-annotated data when evaluated on high-resource languages using the CoNLL-14 benchmark? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.6/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.6/10.

Visit

doi.orgzenodo.org

Tasks

grammar error correction

Tags

performancesyntheticerrordatageneratedzero-shotcross-lingualtransfer

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Scaling Zero-Shot Cross-Lingual Transfer for Synthetic Error Data Generation in Low-Resource LanguagesZero-shot Cross-lingual Transfer Performance of mT5 and Bloom in Synthetic Error Generation for Low-resource LanguagesImpact of Intermediate-Task Training Data Scale on Zero-Shot Cross-Lingual Transfer PerformanceScaling Performance of Zero-Shot Cross-Lingual Retrievers Trained on Synthetic Code-Switched Data on XTREMEGrammatical Error Detection Performance in Low-Resource Languages: Zero-Shot Synthetic vs. Human-Annotated BaselinesComparative Performance of Zero-Shot Synthetic Error Data and Human Annotations on Low-Resource Languages in FLORES-200

Scaling Zero-Shot Cross-Lingual Transfer for Synthetic Error Data Generation in Low-Resource Languages

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Zero-shot Cross-lingual Transfer Performance of mT5 and Bloom in Synthetic Error Generation for Low-resource Languages

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Impact of Intermediate-Task Training Data Scale on Zero-Shot Cross-Lingual Transfer Performance

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Scaling Performance of Zero-Shot Cross-Lingual Retrievers Trained on Synthetic Code-Switched Data on XTREME

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Grammatical Error Detection Performance in Low-Resource Languages: Zero-Shot Synthetic vs. Human-Annotated Baselines

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Comparative Performance of Zero-Shot Synthetic Error Data and Human Annotations on Low-Resource Languages in FLORES-200

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th