Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

Domain:

natural language processing

Record type:

modelpaper
Creator:
KirPet
Publisher:
arXiv
Host:avatar
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.

Visit

doi.org

Tags

Computation and Language (cs.CL)Machine Learning (cs.LG)Robotics (cs.RO)Machine Learning (stat.ML)FOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similar

Nairobi Cannot Plan What It Cannot See: The Case for a Digital Twin of the City's Traffic and Mobility NetworkWhat Can Low Resource Languages Learn From Each Other?Do student teachers see what learners see? – Avoiding instructional dissonance when designing worksheetsWhat a Creole Wants, What a Creole NeedsWhat Do Prompts Reveal About Model Capabilities in Low-Resource Languages?What an Ethics Management Program Cannot Sufficiently Address in an African Context

Nairobi Cannot Plan What It Cannot See: The Case for a Digital Twin of the City's Traffic and Mobility Network

This policy brief argues that Nairobi urgently needs a digital twin of its traffic and mobi

What Can Low Resource Languages Learn From Each Other?

Despite the rapid advancement of Vision-Language Models (VLMs), their linguistic reach remains large

Do student teachers see what learners see? – Avoiding instructional dissonance when designing worksheets

Background: The judicious use of worksheets ought to contribute to the establishment of literacy, wi

What a Creole Wants, What a Creole Needs

In recent years, the natural language processing (NLP) community has given increased attention to th

What Do Prompts Reveal About Model Capabilities in Low-Resource Languages?

Large language models are extremely sensitive to prompt design, a phenomenon which is amplified in m

What an Ethics Management Program Cannot Sufficiently Address in an African Context

Ethics management programs have become a popular first step for organizations to manage ethical risk