Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
ManRacKulRaj
Host:avatar
In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alternative for generating high-quality pretraining data at scale. This is particularly beneficial in low-resource language settings where the benefits of recent LLMs have been unevenly distributed across languages. In this work, we present a systematic study on the generation and evaluation of synthetic multilingual pretraining data for Indic languages, where we construct a large-scale synthetic dataset BhashaKritika, comprising 540B tokens using 5 different techniques for 10 languages. We explore the impact of grounding generation in documents, personas, and topics. We analyze how language choice, both in the prompt instructions and document grounding, affects data quality, and we compare translations of English content with native generation in Indic languages. To support scalable and language-sensitive evaluation, we introduce a modular quality evaluation pipeline that integrates script and language detection, metadata consistency checks, n-gram repetition analysis, and perplexity-based filtering using KenLM models. Our framework enables robust quality control across diverse scripts and linguistic contexts. Empirical results through model runs reveal key trade-offs in generation strategies and highlight best practices for constructing effective multilingual corpora.

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and LanguageArtificial Intelligence

Similar

Multilingual Multi-Label Emotion Classification at Scale with Synthetic DataOn the Utility of Pretraining Language Models on Synthetic DataHow Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality PerspectiveLess is More: Adapting Text Embeddings for Low-Resource Languages with Small Scale Noisy Synthetic DataIndiAnn: An Annotation Platform for Low-Resource Indic LanguagesRethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources

Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data

Emotion classification in multilingual settings remains constrained by the scarcity of annotated dat

On the Utility of Pretraining Language Models on Synthetic Data

Development of pre-trained language models has predominantly relied on large amounts of datasets. Ho

How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality Perspective

Low-resource languages challenge multilingual LLMs due to limited high-quality training data, leadin

Less is More: Adapting Text Embeddings for Low-Resource Languages with Small Scale Noisy Synthetic Data

Low-resource languages (LRLs) often lack high-quality, large-scale datasets for training effective t

IndiAnn: An Annotation Platform for Low-Resource Indic Languages

Linguistic annotation tools that work well for non-Indic languages (e.g. English, German, Spanish, e

Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources

Large Language Models (LLMs) exhibit significant disparities in performance across languages, primar