Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tucano: Advancing Neural Text Generation for Portuguese

Domain:

natural language processing

Record type:

papermodeldataset
Creator:
CorSenFalFat
Host:avatar
Significant advances have been made in natural language processing in recent years. However, our current deep learning approach to language modeling requires substantial resources in terms of data and computation. One of the side effects of this data-hungry paradigm is the current schism between languages, separating those considered high-resource, where most of the development happens and resources are available, and the low-resource ones, which struggle to attain the same level of performance and autonomy. This study aims to introduce a new set of resources to stimulate the future development of neural text generation in Portuguese. In this work, we document the development of GigaVerbo, a concatenation of deduplicated Portuguese text corpora amounting to 200 billion tokens. Via this corpus, we trained a series of decoder-transformers named Tucano. Our models perform equal or superior to other Portuguese and multilingual language models of similar size in several Portuguese benchmarks. The evaluation of our models also reveals that model performance on many currently available benchmarks used by the Portuguese NLP community has little to no correlation with the scaling of token ingestion during training, highlighting the limitations of such evaluations when it comes to the assessment of Portuguese generative language models. All derivatives of our study are openly released on GitHub and Hugging Face. See nkluge-correa.github.io

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similar

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATIONychafiqui/darija-text-generationvictoriapedlar/isizulu-text-generationAmicahh/Amharic-text-generationArabic Large Language Models for Medical Text GenerationGeneration of segmented isiZulu text

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

Machine Translation (MT) poses a significant challenge in developing language corpora for low-resour

ychafiqui/darija-text-generation

victoriapedlar/isizulu-text-generation

Open-Ended Text Generation in isiZulu: Decoding Strategies for a Morphologically Rich Low-Resource L

Amicahh/Amharic-text-generation

# Amharic-text-generation This project implements an Amharic language model based on the GPT archite

Arabic Large Language Models for Medical Text Generation

Efficient hospital management systems (HMS) are critical worldwide to address challenges such as ove

Generation of segmented isiZulu text