Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

GlotStoryBook Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
cis
Host:
Story Books for 180 ISO-639-3 codes. The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset. This dataset consists of 2 subsets: default, which consists of 4 publishers: asp: African Storybook pb: Pratham Books lcb: Little Cree Books lida: LIDA Stories nalibali, which comes from Nal'ibali stories. default: from datasets import load_dataset

Visit

huggingface.co

Tasks

machine translation

Languages

AcholiAfrikaansAlurAmharicAnuakArabic, Tunisian SpokenAringaAtesoBembaBukusu+83

Tags

storybookbookstorylanguage-identificationnalibalimachine-translation

Licenses

cc

Similar

DziriOFN Corpus (Dziri Offensive corpus) v1.0AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-CorpusThe SAWA Corpus: a Parallel Corpus English - SwahiliAfroMAFT Corpus: Language Adaptation Corpus for African languagesBVLAC corpus - Extracted Data Corpus BVLAC - Données extraitesTaxi1500 Corpus

DziriOFN Corpus (Dziri Offensive corpus) v1.0

Dziri refers to the name of the Algerian dialectal Arabic. DziriOFN is a new corpus dedicated to offensive language detection on this under-resourced language. Dziri dialect if known as a complex socio-linguistic situation, where the latter is known by the code-

AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-Corpus

Dataset and code for AAAC: Algerian Arabic Adversarial Corpus for dialect-aware hate speech and prom

The SAWA Corpus: a Parallel Corpus English - Swahili

Research in data-driven methods for Machine Translation has greatly benefited from the increasing availability of parallel corpora. Processing the same text in two different languages yields useful information on how words and phrases are translated from a source l

AfroMAFT Corpus: Language Adaptation Corpus for African languages

Language Adaptation Corpus for 17 African languages, English, French, and Arabic.

BVLAC corpus - Extracted Data Corpus BVLAC - Données extraites

[FR] Dans le cadre du projet SONGES sur la mise en correspondance de données textuelles massives et

Taxi1500 Corpus

This repository contains the raw text data of the Taxi1500-c_v3.0 corpus, without classification lab