Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

On the Role of Bidirectionality in Language Model Pre-Training

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
ArtDu,GoyZet
Hôte:avatar
Prior work on language model pre-training has explored different architectures and learning objectives, but differences in data, hyperparameters and evaluation make a principled comparison difficult. In this work, we focus on bidirectionality as a key factor that differentiates existing approaches, and present a comprehensive study of its role in next token prediction, text infilling, zero-shot priming and fine-tuning. We propose a new framework that generalizes prior approaches, including fully unidirectional models like GPT, fully bidirectional models like BERT, and hybrid models like CM3 and prefix LM. Our framework distinguishes between two notions of bidirectionality (bidirectional context and bidirectional attention) and allows us to control each of them separately. We find that the optimal configuration is largely application-dependent (e.g., bidirectional attention is beneficial for fine-tuning and infilling, but harmful for next token prediction and zero-shot priming). We train models with up to 6.7B parameters, and find differences to remain consistent at scale. While prior work on scaling has focused on left-to-right autoregressive models, our results suggest that this approach comes with some trade-offs, and it might be worthwhile to develop very large bidirectional models. Findings of EMNLP 2022

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similaires

MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili LanguageQalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-trainingFrom Scratch vs Pre-trained: A Dataset Size Analysis for Small-Scale Language Model TrainingTiBERT: Tibetan Pre-trained Language ModelImpact of XLM-R Model Scaling with Adversarial Pre-training on Zero-shot Cross-lingual Performance in XTREME-RDziriBERT: a Pre-trained Language Model for the Algerian Dialect

MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language

Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due

Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-training

Despite remarkable progress in large language models, Urdu-a language spoken by over 230 million peo

From Scratch vs Pre-trained: A Dataset Size Analysis for Small-Scale Language Model Training

This research presents an empirical comparison of from-scratch versus pre-trained language model tra

TiBERT: Tibetan Pre-trained Language Model

The pre-trained language model is trained on large-scale unlabeled text and can achieve state-of-the

Impact of XLM-R Model Scaling with Adversarial Pre-training on Zero-shot Cross-lingual Performance in XTREME-R

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

DziriBERT: a Pre-trained Language Model for the Algerian Dialect

Pre-trained transformers are now the de facto models in Natural Language Processing given their stat