Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Afrikaans Speech Dataset for Whisper Fine-Tuning

Domain:

natural language processing

Record type:

dataset
Creator:
and
Host:
Dataset Card This dataset consists of approximately 56 hours of Afrikaans speech extracted from church sermons, paired with cleaned and aligned transcripts. It is specifically prepared for fine-tuning multilingual ASR models like OpenAI's Whisper (particularly large-v3) on low-resource Afrikaans speech

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Afrikaans

Tags

automatic-speech-recognitionspeechaudioafrikaanslow-resourcemultilingual

Licenses

cc-by-4.0