Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Conversations in Galician: a Large Language Model for an Underrepresented Language

Domain:

natural language processing

Record type:

papermodeldataset
Creator:
BaoPérPar
Host:avatar
The recent proliferation of Large Conversation Language Models has highlighted the economic significance of widespread access to this type of AI technologies in the current information age. Nevertheless, prevailing models have primarily been trained on corpora consisting of documents written in popular languages. The dearth of such cutting-edge tools for low-resource languages further exacerbates their underrepresentation in the current economic landscape, thereby impacting their native speakers. This paper introduces two novel resources designed to enhance Natural Language Processing (NLP) for the Galician language. We present a Galician adaptation of the Alpaca dataset, comprising 52,000 instructions and demonstrations. This dataset proves invaluable for enhancing language models by fine-tuning them to more accurately adhere to provided instructions. Additionally, as a demonstration of the dataset utility, we fine-tuned LLaMA-7B to comprehend and respond in Galician, a language not originally supported by the model, by following the Alpaca format. This work contributes to the research on multilingual models tailored for low-resource settings, a crucial endeavor in ensuring the inclusion of all linguistic communities in the development of Large Language Models. Another noteworthy aspect of this research is the exploration of how knowledge of a closely related language, in this case, Portuguese, can assist in generating coherent text when training resources are scarce. Both the Galician Alpaca dataset and Cabuxa-7B are publicly accessible on our Huggingface Hub, and we have made the source code available to facilitate replication of this experiment and encourage further advancements for underrepresented languages. 5 pages

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and Language

Similar

Command A: An Enterprise-Ready Large Language ModelSinLlama -- A Large Language Model for SinhalaBambaraMLLM: A Unified Multilingual Multimodal Large Language Model for Comprehensive Bambara Language ProcessingBreaking Language Barriers: ASR Model Development for Underrepresented African Languages Using Advanced TechniquesBayLing 2: A Multilingual Large Language Model with Efficient Language AlignmentAn Ubuntu-Guided Large Language Model Framework for Cognitive Behavioral Mental Health Dialogue

Command A: An Enterprise-Ready Large Language Model

In this report we describe the development of Command A, a powerful large language model purpose-bui

SinLlama -- A Large Language Model for Sinhala

Low-resource languages such as Sinhala are often overlooked by open-source Large Language Models (LL

BambaraMLLM: A Unified Multilingual Multimodal Large Language Model for Comprehensive Bambara Language Processing

BambaraMLLM is a unified multilingual multimodal large language model (MMLLM) designed to address th

Breaking Language Barriers: ASR Model Development for Underrepresented African Languages Using Advanced Techniques

Breaking Language Barriers: ASR Model Development for Underrepresented African Languages Using Advanced Techniques

Poster presented at the Deep Learning Indaba 2023 by Joseph Pandeinge Mwatukange

BayLing 2: A Multilingual Large Language Model with Efficient Language Alignment

Large language models (LLMs), with their powerful generative capabilities and vast knowledge, empowe

An Ubuntu-Guided Large Language Model Framework for Cognitive Behavioral Mental Health Dialogue

South Africa's escalating mental health crisis, compounded by limited access to culturally responsiv