Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian Languages

Domain:

natural language processing

Record type:

papermodeldataset
Creator:
CahLovKotPut
Host:avatar
Large language models (LLMs) show remarkable human-like capability in various domains and languages. However, a notable quality gap arises in low-resource languages, e.g., Indonesian indigenous languages, rendering them ineffective and inefficient in such linguistic contexts. To bridge this quality gap, we introduce Cendol, a collection of Indonesian LLMs encompassing both decoder-only and encoder-decoder architectures across a range of model sizes. We highlight Cendol's effectiveness across a diverse array of tasks, attaining 20% improvement, and demonstrate its capability to generalize to unseen tasks and indigenous languages of Indonesia. Furthermore, Cendol models showcase improved human favorability despite their limitations in capturing indigenous knowledge and cultural values in Indonesia. In addition, we discuss the shortcomings of parameter-efficient tunings, such as LoRA, for language adaptation. Alternatively, we propose the usage of vocabulary adaptation to enhance efficiency. Lastly, we evaluate the safety of Cendol and showcase that safety in pre-training in one language such as English is transferable to low-resource languages, such as Indonesian, even without RLHF and safety fine-tuning. Cendol models are released under Apache 2.0 license and will be made publicly available soon

Visit

arxiv.org

Tags

Computation and Language

Similar

NusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language ModelsDhati+: Fine-tuned Large Language Models for Arabic Subjectivity EvaluationA Novel Instruction Tuning Method for Vietnamese Mathematical Reasoning using Trainable Open-Source Large Language ModelsCross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLPLost in Translation: Safety Alignment Failures in Nepali and Code-Switched Variants of Instruction-Tuned Large Language Models184 Evaluating Reasoning-Tuned Large Language Models for Clinical Decision-Making in Spine Surgery

NusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language Models

Large Language Models (LLMs) have demonstrated exceptional promise in translation tasks for high-res

Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces

A Novel Instruction Tuning Method for Vietnamese Mathematical Reasoning using Trainable Open-Source Large Language Models

This study introduces Simple Reasoning with Code (SiRC), a novel instruction fine-tuning method for

Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP

Humanitarian organizations face a critical choice: invest in costly commercial APIs or rely on free

Lost in Translation: Safety Alignment Failures in Nepali and Code-Switched Variants of Instruction-Tuned Large Language Models

Large Language Models (LLMs) are increasingly deployed in multilingual settings, yet safety alignmen

184 Evaluating Reasoning-Tuned Large Language Models for Clinical Decision-Making in Spine Surgery

Abstract Large language models (LLMs) such as OpenAI o1 and DeepSeek-R1 are desi