Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

SIB-Fleurs

Domain:

natural language processing

Record type:

dataset
Creator:
Wue
Host:
SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to. The topics are: Science/Technology Travel Politics Sports Health Entertainment Geography Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*. Dataset creation

Visit

huggingface.co

Tasks

speech processingtext classificationtopic classification

Languages

AfrikaansAkanAmazighAmharicArabic, Egyptian SpokenArabic, Moroccan SpokenArabic, Tunisian SpokenBamanankanBembaChichewa+44

Licenses

cc-by-sa-4.0

Similar

FLEURS SomaliFLEURS-Kobani: Extending the FLEURS Dataset for Northern KurdishFleurs dataset Kinyarwandafeiyixiao/language-identification-fleursvonewman/fleurs-wolof-datasetFLEURS — Ethiopian Languages (v2)

FLEURS Somali

FLEURS Somali is a processed Somali speech dataset derived from the Somali portion of the Google FLE

FLEURS-Kobani: Extending the FLEURS Dataset for Northern Kurdish

FLEURS offers n-way parallel speech for 100+ languages, but Northern Kurdish is not one of them, whi

Fleurs dataset Kinyarwanda

Fleur is a multilingual text and audio dataset. The original dataset was created by Google . The dat

feiyixiao/language-identification-fleurs

Spoken language identification (LID) on 8 typologically diverse languages (Mandarin, Japanese, Engli

vonewman/fleurs-wolof-dataset

FLEURS — Ethiopian Languages (v2)

This dataset is a restructured v2 conversion of the Google FLEURS dataset for two Ethiopian language