Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

IndT5: A Text-to-Text Transformer for 10 Indigenous Languages

Domain:

natural language processing

Record type:

papermodeldataset
Creator:
NagCheAbdCav
Host:avatar
Transformer language models have become fundamental components of natural language processing based pipelines. Although several Transformer models have been introduced to serve many languages, there is a shortage of models pre-trained for low-resource and Indigenous languages. In this work, we introduce IndT5, the first Transformer language model for Indigenous languages. To train IndT5, we build IndCorpus--a new dataset for ten Indigenous languages and Spanish. We also present the application of IndT5 to machine translation by investigating different approaches to translate between Spanish and the Indigenous languages as part of our contribution to the AmericasNLP 2021 Shared Task on Open Machine Translation. IndT5 and IndCorpus are publicly available for research Accepted in AmericasNLP 2021, co-located with NAACL-HLT 2021

Visit

arxiv.org

Tasks

language modelingmachine translation

Tags

Computation and Language

Similar

A transformer-based approach to Nigerian Pidgin text generationTransformer-based coreference resolution modeling for Amharic textTaTA: A Multilingual Table-to-Text Dataset for African LanguagesTaTa: A Multilingual Table-to-Text Dataset for African LanguagesArTST: Arabic Text and Speech TransformerAdversarial Text-to-Speech for low-resource languages

A transformer-based approach to Nigerian Pidgin text generation

Abstract This paper describes the development of a transformer-based text generation model for Nige

Transformer-based coreference resolution modeling for Amharic text

Abstract Coreference resolution is the task of identifying

TaTA: A Multilingual Table-to-Text Dataset for African Languages

TaTA (Table-to-Text in African languages) is the first large multilingual table-to-text datasets with a focus on African languages. The dataset is parallel and covers nine languages, eight of which are spoken in Africa: Arabic, English, French, Hausa, Igbo, Portugu

TaTa: A Multilingual Table-to-Text Dataset for African Languages

Existing data-to-text generation datasets are mostly limited to English. To address this lack of data, we create Table-to-Text in African languages (TaTa), the first large multilingual table-to-text dataset with a focus on African languages. We created TaTa by tran

ArTST: Arabic Text and Speech Transformer

We present ArTST, a pre-trained Arabic text and speech transformer for supporting open-source speech

Adversarial Text-to-Speech for low-resource languages

Improving the adversarial TTS models for low-resource languages by utilizing the high-frequency similarities between the different languages.