Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
InoAlhBaiBou
Hôte:avatar
In this paper, we explore the effects of language variants, data sizes, and fine-tuning task types in Arabic pre-trained language models. To do so, we build three pre-trained language models across three variants of Arabic: Modern Standard Arabic (MSA), dialectal Arabic, and classical Arabic, in addition to a fourth language model which is pre-trained on a mix of the three. We also examine the importance of pre-training data size by building additional models that are pre-trained on a scaled-down set of the MSA variant. We compare our different models to each other, as well as to eight publicly available models by fine-tuning them on five NLP tasks spanning 12 datasets. Our results suggest that the variant proximity of pre-training data to fine-tuning data is more important than the pre-training data size. We exploit this insight in defining an optimized system selection model for the studied tasks. Accepted to WANLP 2021

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and Language

Similaires

Arabic Extractive Summarization Using Pre-Trained ModelsMorphosyntactic Tagging with Pre-trained Language Models for Arabic and its DialectsIMPROVING ARABIC TEXT SUMMARIZATION USING ADVANCED PRE-TRAINED MODELSAbusive and Hate speech Classification in Arabic Text Using Pre-trained Language Models and Data AugmentationHow Linguistically Fair Are Multilingual Pre-Trained Language Models?LAWDR: Language-Agnostic Weighted Document Representations from Pre-trained Models

Arabic Extractive Summarization Using Pre-Trained Models

Automatic Text Summarization (ATS) is a crucial area of study in Natural Language Processing (NLP) d

Morphosyntactic Tagging with Pre-trained Language Models for Arabic and its Dialects

We present state-of-the-art results on morphosyntactic tagging across different varieties of Arabic

IMPROVING ARABIC TEXT SUMMARIZATION USING ADVANCED PRE-TRAINED MODELS

The exponential growth of online content has made the task of locating specific information increasi

Abusive and Hate speech Classification in Arabic Text Using Pre-trained Language Models and Data Augmentation

Hateful content on social media is a worldwide problem that adversely affects not just the targeted

How Linguistically Fair Are Multilingual Pre-Trained Language Models?

Massively multilingual pre-trained language models, such as mBERT and XLM-RoBERTa, have received sig

LAWDR: Language-Agnostic Weighted Document Representations from Pre-trained Models

Cross-lingual document representations enable language understanding in multilingual contexts and al