Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

NADiA: News Articles Dataset in Arabic for Multi-Label Text Categorization

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Eln
Éditeur:
Al-Ein
Éditeur:
Men
Hôte:avatar
NADiA Dataset is the largest, to the best of our knowledge, source for Arabic textual data that can be used in any NLP related task such as text classification. We chose the abbreviation NADiA as it is a common Arabic name. The data was collected by scraping ‘SkyNewsArabia’ and ‘Masrawy’ news websites using Python scripts that are fine-tuned for each website. SkyNewsArabia will be referred to as NADiA1, while the latter would be NADiA2. NADiA1 is a big dataset containing 37,445 files, while NADiA2 is a huge dataset that contains 678,563 files. However, after filtering and cleaning we reduced the numbers to 35,416 and 451,230 for NADiA 1 and 2, respectively. NADiA1 consists of the following categories (24, displayed in English for easy referencing): News, North Africa, Levant, Middle East, The Americas, Research, Finance & Economy, War & Terrorism, Gulf, Europe, Political Figures, Iran, Technology, Russia, Sports, Tennis, Football, English League, Arabian Sports, Spanish League, Health, East Asia, Environment, Other Countries NADiA2 consists of the following categories (28, displayed in English for easy referencing): Politics, Middle East, Asia, Africa, United States, Europe, Other Countries, Leaders, Sports, Arabian Sports, Football Clubs, Spanish League, Egyptian League, Finance, Arts, Cinema & TV, Fashion, Health, Pregnancy & Delivery, Cancer, Obesity, Social Media, Technology, Religion, Islamic, Fatawa, Worship, Prophet Biography

Visit

doi.orgdata.mendeley.com

Tasks

topic classificationtext classification

Tags

Natural Language ProcessingMachine LearningClassification SystemInformation ClassificationArabic LanguageCategorizationText ProcessingDeep Learning

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

AraTox: A Multi-Dialect, Multi-Label Arabic Dataset for Toxicity DetectionAFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING DEEP LEARNING APPROACHAUTOMATIC AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING NEURAL NETWORK APPROACHDataset for "Enhancing multi-label emotion analysis in Indonesian social media with emoji-aware text representations"Multi‐dimensional long short‐term memory networks for artificial Arabic text recognition in news videoApplication of Data Mining Classification Algorithms for Afaan Oromo Media Text News Categorization

AraTox: A Multi-Dialect, Multi-Label Arabic Dataset for Toxicity Detection

AraTox is a multi-dialect, multi-label Arabic dataset for toxicity detection. It contains annotated

AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING DEEP LEARNING APPROACH

The development of the internet has made Afaan Oromo's writings widely available both offline and on

AUTOMATIC AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING NEURAL NETWORK APPROACH

Major Advisor: - GetachewMamo(PHD) The classification of natural language texts has gained a growin

Dataset for "Enhancing multi-label emotion analysis in Indonesian social media with emoji-aware text representations"

Creator Amalia Amalia1*, Maya Silvi Lydia1, Rahmi Putri Rangkuti2, Farhan Purwanto Marulitua3, Fikr

Multi‐dimensional long short‐term memory networks for artificial Arabic text recognition in news video

This study presents a novel approach for Arabic video text recognition based on recurrent neural net

Application of Data Mining Classification Algorithms for Afaan Oromo Media Text News Categorization