Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

On the Surprising Efficacy of Distillation as an Alternative to Pre-Training Small Models

Type de record:

paper
Créateur:
FarChe
Hôte:avatar
In this paper, we propose that small models may not need to absorb the cost of pre-training to reap its benefits. Instead, they can capitalize on the astonishing results achieved by modern, enormous models to a surprising degree. We observe that, when distilled on a task from a pre-trained teacher model, a small model can achieve or surpass the performance it would achieve if it was pre-trained then finetuned on that task. To allow this phenomenon to be easily leveraged, we establish a connection reducing knowledge distillation to modern contrastive learning, opening two doors: (1) vastly different model architecture pairings can work for the distillation, and (2) most contrastive learning algorithms rooted in the theory of Noise Contrastive Estimation can be easily applied and used. We demonstrate this paradigm using pre-trained teacher models from open-source model hubs, Transformer and convolution based model combinations, and a novel distillation algorithm that massages the Alignment/Uniformity perspective of contrastive learning by Wang & Isola (2020) into a distillation objective. We choose this flavor of contrastive learning due to its low computational cost, an overarching theme of this work. We also observe that this phenomenon tends not to occur if the task is data-limited. However, this can be alleviated by leveraging yet another scale-inspired development: large, pre-trained generative models for dataset augmentation. Again, we use an open-source model, and our rudimentary prompts are sufficient to boost the small model`s performance. Thus, we highlight a training method for small models that is up to 94% faster than the standard pre-training paradigm without sacrificing performance. For practitioners discouraged from fully utilizing modern foundation datasets for their small models due to the prohibitive scale, we believe our work keeps that door open. ICLR 2024. 5th Workshop on Practical ML for Low Resource Settings (PML4LRS). Code can be found at github.com

Visit

arxiv.org

Tags

Machine LearningArtificial Intelligence

Similaires

Stable Distillation: Regularizing Continued Pre-training for Low-Resource Automatic Speech RecognitionSmall Languages, Big Models: A Study of Continual Training on Languages of NorwayPerformance Variation in Multilingual Pre-trained Language Models with Optimal Transport Distillation for AdversarialDomain Specific Specialization in Low-Resource Settings: The Efficacy of Offline Response-Based Knowledge Distillation in Large Language ModelsLecture 3 : « Language contact as an alternative to assumed genealogical relationships »Small language models efficacy prototyped for Oromo word sense disambiguation

Stable Distillation: Regularizing Continued Pre-training for Low-Resource Automatic Speech Recognition

Continued self-supervised (SSL) pre-training for adapting existing SSL models to the target domain h

Small Languages, Big Models: A Study of Continual Training on Languages of Norway

Training large language models requires vast amounts of data, posing a challenge for less widely spo

Performance Variation in Multilingual Pre-trained Language Models with Optimal Transport Distillation for Adversarial

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Domain Specific Specialization in Low-Resource Settings: The Efficacy of Offline Response-Based Knowledge Distillation in Large Language Models

Large Language Models (LLMs) excel in general tasks but often struggle with hallucinations when hand

Lecture 3 : « Language contact as an alternative to assumed genealogical relationships »

« Areal linguistics in Africa before a new approach to its genealogical language classification », L

Small language models efficacy prototyped for Oromo word sense disambiguation

Abstract Word Sense Disambiguation (WSD) is a fundamental i