Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Aya Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
You
Host:
The Aya Dataset is a multilingual instruction fine-tuning dataset curated by an open-science community via Aya Annotation Platform from Cohere Labs. The dataset contains a total of 204k human-annotated prompt-completion pairs along with the demographics data of the annotators. This dataset can be used to train, finetune, and evaluate multilingual LLMs. Curated by: Contributors of Aya Open Science Intiative.

Visit

huggingface.co

Languages

AmharicArabic, Egyptian SpokenArabic, Moroccan SpokenChichewaHausaIgboMalagasyShonaSomaliSotho, Northern+5

Licenses

apache-2.0

Similar

Aya DatasetThe Aya DatasetAya African Alpaca Style DatasetTiny-Aya-Base Blind Spots Evaluation DatasetAya Dataset: An Open-Access Collection for Multilingual Instruction TuningCohereLabsCommunity/afri-aya

Aya Dataset

The Aya Dataset is a multilingual instruction fine-tuning dataset curated by an open-science communi

The Aya Dataset

The Aya Dataset is a multilingual instruction fine-tuning dataset curated by an open-science community via Aya Annotation Platform from Cohere For AI. The dataset contains a total of 204k human-annotated prompt-completion pairs along with the demographic data of th

Aya African Alpaca Style Dataset

This dataset card aims to be a base template for new datasets. It has been generated using this raw

Tiny-Aya-Base Blind Spots Evaluation Dataset

CohereLabs/tiny-aya-base Architecture: Cohere2 (Cohere Command R family) Parameters: ~3B Type: Raw

Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of tasks that enables a large la

CohereLabsCommunity/afri-aya

Giving Sight to African LLMs Afri-Aya is a community-curated multilingual image dataset covering 13