Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

South African Broadcast News (SABN) Corpus

Domain:

natural language processing

Record type:

dataset
Publisher:
Stellenbosch UniversityCSIR
Host:avatar
The corpus consists of approximately 20 hours of audio recordings from one of the country's main radio news channels, SAFM. Bulletins were broadcast between 1996 and 2006 and are a mix of news-reader speech, interviews, and crossings to reporters

Visit

hdl.handle.net

Tags

broadcast news transcriptionSouth African Englishaccents of Englishunder-resourced languages

Similar

African News CorpusTigre Broadcast Speech CorpusPortuguese Variety Identification on Broadcast NewsOverlay Text Extraction From TV News BroadcastLanguage and Variety Verification on Broadcast News for PortugueseSwahili News Corpus

African News Corpus

This consist of a monolingual news corpus for 19 languages from various sources like VOA, B

Tigre Broadcast Speech Corpus

A large-scale, open-source speech dataset for the Tigre language (ISO 639-3: tig), developed to supp

Portuguese Variety Identification on Broadcast News

International audience This paper describes an accent identification system for Portu

Overlay Text Extraction From TV News Broadcast

The text data present in overlaid bands convey brief descriptions of news events in broadcast videos

Language and Variety Verification on Broadcast News for Portuguese

International audience This paper describes a language/accent verification system for

Swahili News Corpus

Language modeling, topic classification, AI training for Swahili NLP, digital literacy tools Notes