Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Moroccan News Dataset and Source Code for Arabic and French Topic Detection

Domain:

natural language processing

Record type:

datasetsoftware
Creator:
Hal
Publisher:
Men
Host:avatar
This dataset contains news articles collected from 37 Moroccan press sources (21 Arabic and 16 French) via RSS feeds over a period of 8 months, from June 14, 2025 to February 14, 2026. The dataset comprises 729 Arabic files and 729 French files, corresponding to one file per system update (3 updates per day). After deduplication, the dataset contains 115145 unique Arabic articles and 55239 French articles. Each file is stored in Python Pickle format and contains a list of dictionaries, where each dictionary represents a news article with the following fields : title, content, source, url, and date. A Python script for reading and preprocessing the data is provided alongside the dataset. This dataset was created to support a study on automated news topic detection. Along side the dataset, we provide the source code of the of the proposed monitoring pipeline : data collection, preprocessing, transformer-based (BERT/AraBERT) contextual embedding, topic detection using Graph Community Analysis, and treemap and graph visualizations. To simplify deployment, the code comes with transformer models and NLP tools, allowing for immediate use on the provided data. Detailed instructions on how to set up the environment and run the topic detection pipeline are provided in the README file.

Visit

doi.orgdata.mendeley.com

Tasks

topic classificationtext classification

Tags

Artificial IntelligenceInformation RetrievalNatural Language Processing

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution Non Commercial 4.0 Internationalhttps://creativecommons.org/licenses/by-nc/4.0/legalcode

Similar

Code Switching Between Moroccan Arabic and FrenchEfficient Topic Detection System for Online Arabic News Between Arabic and French Lies the Dialect: Moroccan Code-Weaving on FacebookCAFE: Algerian Arabic, French, and English Code-Switched Conversational Speech Dataset ANTC — African News Topic Classification DatasetMoroccan Darija Code-Switched Corpus (Moroccan Arabic)

Code Switching Between Moroccan Arabic and French

Efficient Topic Detection System for Online Arabic News

Between Arabic and French Lies the Dialect: Moroccan Code-Weaving on Facebook

This thesis examines code-switching in Morocco. Specifically, it looks at the Morocco's linguistic h

CAFE: Algerian Arabic, French, and English Code-Switched Conversational Speech Dataset

The CAFE dataset (Code-switching Algerian French English) addresses the lack of publicly available r

ANTC — African News Topic Classification Dataset

We created a novel dataset, ANTC — African News Topic Classification for 4 African languages. We obtained data from three different news sources: VOA, BBC6 and isolezwe7 . From the VOA data we created datasets for Lingala and Somali. We obtained the topics from dat

Moroccan Darija Code-Switched Corpus (Moroccan Arabic)

This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per