Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Corpus of African Digital News from 600 Websites Formatted for Text Mining / Computational Text Analysis

Domaine:

natural language processing

Type de record:

dataset
Créateur:
MadLinPer
Éditeur:
Mad
Éditeur:
Tex
Hôte:avatar
This dataset includes a corpus 200,000+ news articles published by 600 African news organizations between December 4, 2020 and January 3, 2021. The texts have been pre-processed (punctuation and English stopwords have been removed, features have been lowercased, lemmatized and POS-tagged) and stored in commonly used formats for text mining/computational text analysis. Users are advised to read the documentation for an explanation of the data collection process.

This dataset includes the following items:
  • 31 tables (one per day) of lowercased and lemmatized tokens with the following additional variables: POS tags, document id, sentence id, token id and publication date (stored as a tibble).
  • A single document-feature matrix (DFM) with raw counts of feature frequencies in each news article (stored as a quanteda dfm object). The DFM comes with the following metadata for each document: date of publication and source URL.
  • A metadata table with the following fields: document id, publication date, source url, news source and country of the news source.
  • A list of sources included in the course grouped by country name.
All items are stored in formats readable in R. The documentation provides instructions on how to load the RDS files to R.

If you decide to use the data for your own project, please do cite it using the information above. If you identify errors or missing sources, please contact us so that these can be addressed. R, 4.2

Visit

doi.orgdataverse.tdl.org

Tags

Arts and HumanitiesSocial SciencesnewsAfricadigital newstext miningcomputational

Licenses

info:eu-repo/semantics/openAccessCreative Commons Zero v1.0 Universalhttps://creativecommons.org/publicdomain/zero/1.0/legalcode

Similaires

Using Computational Text Analysis Tools to Study African Online News ContentA Computational Mapping of Online News Deserts on African News WebsitesHtgotcode/Twitter-ZA-News-Data-Collection-and-Text-MiningAfrican Multilingual Text CorpusApplication of Data Mining Classification Algorithms for Afaan Oromo Media Text News CategorizationWeb-Scraped Nigerian Pidgin English Text Dataset from Digital News Platforms

Using Computational Text Analysis Tools to Study African Online News Content

A Computational Mapping of Online News Deserts on African News Websites

To date, the study of news deserts, geographic spaces lacking local news and information

Htgotcode/Twitter-ZA-News-Data-Collection-and-Text-Mining

South African news Twitter accounts will have their timeline analyzed on how their sentiment & topic

African Multilingual Text Corpus

NLP model training, translation Notes / challenges: Unequal language representation

Application of Data Mining Classification Algorithms for Afaan Oromo Media Text News Categorization

Web-Scraped Nigerian Pidgin English Text Dataset from Digital News Platforms

This dataset consists of Nigerian Pidgin English text collected through web scraping of multiple Nig