Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

kreasof-ai/bemba-speech-csikasote

Domain:

natural language processing

Record type:

dataset
Creator:
kre
Host:
This is speech dataset of Bemba language. This dataset was acquired from (BembaSpeech)[github.com]. BembaSpeech is the speech recognition corpus in Bemba [1]. DatasetDict({ train: Dataset({ features: ['audio', 'sentence'], num_rows: 12421 }) dev: Dataset({ features: ['audio', 'sentence'], num_rows: 1700 }) test: Dataset({

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

AfrikaansBemba

Similar

csikasote/xls-r-bemba-expskreasof-ai/bigc-bem-engkreasof-ai/flores200-eng-bemkreasof-ai/tatoeba-eng-bem-backtranslationcsikasote/iwslt_2026_bem_eng_testcsikasote/bigc

csikasote/xls-r-bemba-exps

Scripts to finetune the multilingual Facebook`s wav2vec2 xls-r models on BembaSpeech as adapted from

kreasof-ai/bigc-bem-eng

This is dataset of speech translation task for Bemba-to-English Language. This dataset is acquired f

kreasof-ai/flores200-eng-bem

This is Bemba-to-English dataset for machine translation task. This dataset is a customized version

kreasof-ai/tatoeba-eng-bem-backtranslation

This dataset contains Bemba-to-English sentences which is intended to machine translation task. This

csikasote/iwslt_2026_bem_eng_test

Test data for the IWSLT 2025 Workshop Low-Resource Shared Task for Bemba/English language pair ###

csikasote/bigc

This repository contains the data resources for the LacunaFund supported project, Multimodal dataset