Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tigre Wikipedia Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Bei
Host:
This repository houses the Tigre Wikipedia Corpus, a foundational linguistic resource containing all non-template articles from tig.wikipedia.org. Tigre is an under-resourced South Semitic language within the Afro-Asiatic family.

Visit

huggingface.co

Languages

Tigré

Tags

wikipediatigre languagetigcorpuslow-resource

Licenses

cc-by-sa-4.0

Similar

Tigre Broadcast Speech CorpusKalebu/wikipedia-swahili-corpusrashiedomar/somali-wikipedia-corpusTigreThe Etymology of "Tigre" in the Tigre LanguageTigre language

Tigre Broadcast Speech Corpus

A large-scale, open-source speech dataset for the Tigre language (ISO 639-3: tig), developed to supp

Kalebu/wikipedia-swahili-corpus

A Swahili corpus made from Swahili Wikipedia articles # wikipedia-swahili-corpus A Swahili corpus m

rashiedomar/somali-wikipedia-corpus

Cleaned Somali Wikipedia corpus (~9,500 articles) for NLP, LLM training, and linguistic research #

Tigre

The Etymology of "Tigre" in the Tigre Language

Tigre language

This repository introduces the Monolingual Text component of the Tigre language resource collection.