Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

thirtyninetythree/kikuyu_monolingual_sentences

Domain:

natural language processing

Record type:

dataset
Creator:
thi
Host:
This dataset is a combined monolingual corpus of Kikuyu (Gikuyu) text, created by extracting the Kikuyu language portions from 38 existing parallel datasets on the Hugging Face Hub. The raw combined size before deduplication was approximately 2.87 million entries. The final unique deduplicated dataset contains 118,887 entries (as per num_examples in dataset_info). Original datasets combined (all from the michsethowusu organization):

Visit

huggingface.co

Tasks

language modeling

Languages

Gikuyu

Similar

thirtyninetythree/kikuyu-bpe-tokenizerthirtyninetythree/TinyLlama-1.1B-Kikuyu-LoRAthirtyninetythree/kikuyu-Llama-3.2-3B-lorathirtyninetythree/kikuyu-lora-llama-3.2-3b-instruct

thirtyninetythree/kikuyu-bpe-tokenizer

thirtyninetythree/TinyLlama-1.1B-Kikuyu-LoRA

thirtyninetythree/kikuyu-Llama-3.2-3B-lora

thirtyninetythree/kikuyu-lora-llama-3.2-3b-instruct