The Zarma_POS dataset is a part-of-speech (POS) tagged corpus for the Zarma language, derived from the 27Group/Feriji dataset's fr_dje_corpus subset.
Each entry in the dataset contains:
text: The original Zarma sentence.
tokens: A list of tokenized words and punctuation.
tags: A list of POS tags corresponding to each token (e.g., NOUN, VERB, PUNCT).
{
"text": "Waybora di alboro.",