Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

From Scratch vs Pre-trained: A Dataset Size Analysis for Small-Scale Language Model Training

Domain:

natural language processing

Record type:

paper
Creator:
CHA
Publisher:
Zenodo
Host:avatar
This research presents an empirical comparison of from-scratch versus pre-trained language model training strategies across small dataset sizes (1MB-20MB). The study reveals a critical dataset size threshold effect at 5-10MB where optimal training strategies diverge. Key findings include: from-scratch models achieving superior generalization on datasets <5MB through memorization mechanisms (PPL=1.0), while pre-trained models excel on datasets >10MB through genuine pattern learning. The work provides adaptive hyperparameter strategies and practical guidelines for model selection in resource-constrained scenarios. Implications extend to specialized domains, low-resource languages, and situations where gigabyte-scale datasets are unavailable.

Visit

doi.orgzenodo.org

Tasks

language modeling

Tags

Transfer LearningSmall Data LearningDataset Size AnalysisFrom Scratch TrainingPre-trained ModelsLanguage ModelsMachine LearningDeep LearningMemorization vs GeneralizationOverfitting+21

Licenses

Creative Commons Attribution Non Commercial No Derivatives 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-nd/4.0/legalcode© 2025 Théo (RDTvlokip) — All rights reserved under CC BY-NC-ND 4.0http://rightsstatements.org/vocab/InC/1.0/