Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TCNSpeech: A Community-Curated Speech Corpus for Sermons

Domain:

natural language processing

Record type:

paper
In this work we present TCNSpeech, a community-curated multispeaker sermon corpus for speech recognition tasks. It contains a total of 24 hours of English audio data recording, chunked and transcribed. The context of the dataset is domain-specific for sermons in Nigerian English accent and a use case for community data curation. The dataset will be made publicly available.

Visit

openreview.net

Tasks

automatic speech recognitionspeech processing

Languages

Pidgin, Nigerian

Tags

africanlp2sermons

Similar

A Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech SynthesisThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionDesign of a Tigrinya Language Speech Corpus for Speech RecognitionOMAN-SPEECH: A Multi-Layer Annotated Speech Corpus for Omani Arabic DialectsKurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-Speech

A Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech Synthesis

This dataset contains curated and preprocessed speech recordings in Luganda and Kiswahili for use in

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda. The corp

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the Un

Design of a Tigrinya Language Speech Corpus for Speech Recognition

In this paper, we describe the first Tigrinya Languages speech corpora designed and development for speech recognition purposes. Tigrinya, often written as Tigrigna (ትግርኛ) /tɪˈɡrinjə/ belongs to the Semitic branch of the Afro-Asiatic languages where it shows the ch

OMAN-SPEECH: A Multi-Layer Annotated Speech Corpus for Omani Arabic Dialects

Automatic Speech Recognition (ASR) has achieved strong performance in high-resource languages; howev

KurFemTTS: A Large-Scale Kurdish Female Speech Corpus for Text-to-Speech

KurFemTTS is a collaborative initiative between the University of Kurdistan Hewlêr (UKH) and Kurdish