Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Kanuri Books Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
CLEAR Global
Hôte:
A text corpus of 10,281 randomized sentences (90,706 words) extracted from books by Kanuri authors Dr. Baba Kura Alkali Gazali, Lawan Dalama, Kaka Gana Abba, and Lawan Hassan. The corpus includes both original and normalized (lowercased, punctuation-removed) versions. It was compiled by CLEAR Global (formerly Translators without Borders) for the creation of open-source language technology. These sentences were also recorded by multiple speakers to make a speech corpus published within TWB Voice.

Visit

mozilladatacollective.com

Languages

KanembuKanuri, MangaKanuri, Yerwa

Tags

mdcmozilla data collectiveLMTXT

Licenses

Creative Commons Attribution 4.0 International (CC-BY-4.0)

Similaires

CLEAR-Global/kanuri-books-corpus

CLEAR-Global/kanuri-books-corpus

Randomized sentences from books collected from Kanuri authors: Dr. Baba Kura Alkali Gazali, Lawan Da