Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

South African Broadcast News (SABN) Corpus

Domaine:

natural language processing

Type de record:

dataset
Éditeur:
Stellenbosch UniversityCSIR
Hôte:avatar
The corpus consists of approximately 20 hours of audio recordings from one of the country's main radio news channels, SAFM. Bulletins were broadcast between 1996 and 2006 and are a mix of news-reader speech, interviews, and crossings to reporters

Visit

hdl.handle.net

Tags

broadcast news transcriptionSouth African Englishaccents of Englishunder-resourced languages

Similaires

African News CorpusTigre Broadcast Speech CorpusPortuguese Variety Identification on Broadcast NewsOverlay Text Extraction From TV News BroadcastLanguage and Variety Verification on Broadcast News for PortugueseSwahili News Corpus

African News Corpus

This consist of a monolingual news corpus for 19 languages from various sources like VOA, B

Tigre Broadcast Speech Corpus

A large-scale, open-source speech dataset for the Tigre language (ISO 639-3: tig), developed to supp

Portuguese Variety Identification on Broadcast News

International audience This paper describes an accent identification system for Portu

Overlay Text Extraction From TV News Broadcast

The text data present in overlaid bands convey brief descriptions of news events in broadcast videos

Language and Variety Verification on Broadcast News for Portuguese

International audience This paper describes a language/accent verification system for

Swahili News Corpus

Language modeling, topic classification, AI training for Swahili NLP, digital literacy tools Notes