Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

On the Utility of Pretraining Language Models on Synthetic Data

Domain:

natural language processing
Creator:
AssAbdAlcKwo
Publisher:
Und
Host:avatar
Development of pre-trained language models has predominantly relied on large amounts of datasets. However, this dependence on abundant data has limited the applicability of these models in low-resource settings. In this work, we investigate the utility of exploiting synthetic datasets acquired from different sources to pre-train language models for Arabic. Namely, we leverage data derived based on four different methods: optical character recognition (OCR), automatic speech recognition (ASR), machine translation (MT), and generative language models. We use these datasets to pre-train models in three different architectures: encoder-only (BERTtextsubscript{Base}), encoder-decoder (T5), and decoder-only (GPT-2). We test the capabilities of resulting models on Arabic natural language understanding (NLU) tasks using the ORCA benchmark. Our results show that utilizing synthetic data can achieve performance comparable to, or even surpassing, those trained on gold data. For example, our model based on a GPT-2 architecture trained on a combined synthetic dataset surpasses the baseline model ARBERTtextsubscript{v2}. Overall, our models pre-trained on synthetic data demonstrate robust performance across various tasks. This highlights the potential of synthetic datasets in augmenting language model training in low-resource settings.

Visit

doi.orgunderline.io

Tasks

language modeling

Tags

Computational LinguisticsNatural Language ProcessingLinguisticsFOS: Languages and literature

Similar

One Script Instead of Hundreds? On Pretraining Romanized Encoder Language ModelsHow Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality PerspectiveData Caricatures: On the Representation of African American Language in Pretraining CorporaBhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languagesmohammad-gh009/Small-language-models-on-clinical-data-extractionEffect of African Language Pretraining on XTREME-R Cross-Lingual Transfer Performance

One Script Instead of Hundreds? On Pretraining Romanized Encoder Language Models

Exposing latent lexical overlap, script romanization has emerged as an effective strategy for improv

How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality Perspective

Low-resource languages challenge multilingual LLMs due to limited high-quality training data, leadin

Data Caricatures: On the Representation of African American Language in Pretraining Corpora

With a combination of quantitative experiments, human judgments, and qualitative analyses, we evalua

BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages

In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alte

mohammad-gh009/Small-language-models-on-clinical-data-extraction

This repository contains the implementation and supporting code for the research study “Small Langua

Effect of African Language Pretraining on XTREME-R Cross-Lingual Transfer Performance

Recent advances in training multilingual language models on large datasets seem to have shown promis