Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Impact of Pretraining Data Volume on Zero-Shot Cross-Lingual Code Generation Performance

Domain:

natural language processing

Record type:

paper
Creator:
Ass
Publisher:
Zenodo
Host:avatar
Zero-shot cross-lingual transfer is a central task in multilingual NLP, allowing models trained in languages with more sufficient training resources to generalize to other low-resource languages. Earlier efforts on this task use parallel corpora, bilingual dictionaries, or other annotated alignment data to improve cross-lingual transferability, which are typically expensive to obtain. In this paper, we propose a simple yet effective method, SALT, to improve the zero-shot cross-lingual transfer of the multilingual pretrained language models without the help of such external data. By incorporati Research goal: To what extent does increasing pretraining data volume for low-resource languages improve zero-shot cross-lingual transfer performance on code generation tasks compared to architecture scaling? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.5/10.

Visit

doi.org

Tags

extentincreasingpretrainingdatavolumelow-resourcelanguagesimprove

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode