Logo Lanfrica

Reliable Generative Data Augmentation for Arabic Text Classification: A Length-Aware Data-Centric Approach

Domaine:

natural language processing

Type de record:

paper
Créateur:
Man
Éditeur:
Elsevier BV
Hôte:
Data augmentation has become a critical strategy for improving text classification performance, particularly in low-resource settings such as Arabic Natural Language Processing (NLP), where annotated data remain limited and highly variable across dialects and domains. However, existing augmentation methods often suffer from semantic drift, label inconsistency, and limited adaptability to text length and context. This paper proposes a length-aware generative data augmentation framework that integrates transformer-based language models with an adaptive semantic filtering mechanism to enhance data quality and label preservation. The framework dynamically applies context-independent or context-aware conditioning based on input length and employs sentence-level embeddings to filter generated samples using a similarity-based adaptive threshold. This design enables controlled diversity while minimizing noise and label corruption. Extensive experiments are conducted on both short-text sentiment analysis and longtext news classification datasets, covering Modern Standard Arabic and dialectal Arabic. The proposed approach consistently outperforms baseline and existing augmentation techniques, achieving improvements of up to 1.90% in Macro F1-score. Statistical significance analysis confirms that the observed gains are robust across multiple runs (𝑝 < 0.05), while qualitative analysis demonstrates improved semantic consistency and reduced error patterns such as negation loss. The results highlight the importance of combining generative modeling with semantic filtering for reliable data augmentation. The proposed framework provides a scalable and effective solution for enhancing Arabic text classification performance in low-resource scenarios and establishes a strong foundation for future research in generative augmentation and data-centric NLP