Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Urdu Alpaca

Domain:

natural language processing

Record type:

dataset
Creator:
PRO
Host:
This dataset is an Urdu language translation of the Stanford Alpaca instruction tuning dataset, filtered to 28,910 rows. Each entry contains an instruction, optional input, and output translated into Urdu, covering a broad range of task types including question answering, summarization, creative writing, classification, and reasoning. The dataset is designed to support instruction-tuning of Urdu language large language models, extending Alpaca style instruction data to a low resource language.

Visit

mozilladatacollective.com

Tasks

language modelingmachine translation

Tags

mdcmozilla data collectiveNLGParquet

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)

Similar

Sequence to Sequence Networks for Roman-Urdu to Urdu TransliterationSwahili AlpacaAlpaca SwahiliKinyarwanda alpaca-52kMalagasy alpaca-52kZulu alpaca-52k

Sequence to Sequence Networks for Roman-Urdu to Urdu Transliteration

Neural Machine Translation models have replaced the conventional phrase based statistical translatio

Swahili Alpaca

Alpaca dataset for instruction fine-tuning in Swahili. ### Maelekezo:\n{instruction} ### Agizo:\n{i

Alpaca Swahili

Alpaca Swahili is a dataset of 52,000 instructions and translated from the latest Alpaca Dataset. Re

Kinyarwanda alpaca-52k

This repository contains the dataset used for the TaCo paper. The dataset follows the style outlined

Malagasy alpaca-52k

This repository contains the dataset used for the TaCo paper. The dataset follows the style outlined

Zulu alpaca-52k

This repository contains the dataset used for the TaCo paper. The dataset follows the style outlined