Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kanuri Books Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
CLEAR Global
Host:
A text corpus of 10,281 randomized sentences (90,706 words) extracted from books by Kanuri authors Dr. Baba Kura Alkali Gazali, Lawan Dalama, Kaka Gana Abba, and Lawan Hassan. The corpus includes both original and normalized (lowercased, punctuation-removed) versions. It was compiled by CLEAR Global (formerly Translators without Borders) for the creation of open-source language technology. These sentences were also recorded by multiple speakers to make a speech corpus published within TWB Voice.

Visit

mozilladatacollective.com

Languages

KanembuKanuri, MangaKanuri, Yerwa

Tags

mdcmozilla data collectiveLMTXT

Licenses

Creative Commons Attribution 4.0 International (CC-BY-4.0)

Similar

CLEAR-Global/kanuri-books-corpus

CLEAR-Global/kanuri-books-corpus

Randomized sentences from books collected from Kanuri authors: Dr. Baba Kura Alkali Gazali, Lawan Da