Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Nigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer

Domain:

socioeconomic

Record type:

dataset
Creator:
BasFak
Publisher:
arXiv
Host:avatar
There is relatively little, public, and model-ready data on industrial machinery for African economies. This makes it hard to do quantitative analysis or to train language models on numeric tasks grounded in that setting. We release two things to help with part of this problem. The first is the Nigeria Machinery Usage and Failures Dataset: 89 machine-level records across 28 indicators, covering Nigeria's manufacturing and oil and gas sectors from 2006 to 2025. Every record names a public source and is decoded by a codebook. The second is a method for building chain-of-thought (CoT) reasoning examples from these sparse numeric values. The result is 94 prompt, completion, and reasoning-trace rows. In every row, the prompt names the real indicator, subsector, year, and source of the record it comes from. The data adaptation work was carried out by Adaption Labs. Along the way we describe a problem that is common when language models are used to build datasets. The prompts can match the real numbers while saying nothing about the real domain. We show that fixing this raises the share of domain-grounded prompts from 1 out of 78 in an earlier release to 94 out of 94, and that every retrieval answer now matches its source value (84 out of 84). We release the data, the reasoning layer, and a per-row provenance file under CC-BY-4.0. We are clear about the limits. With 89 records and 17 indicators that have only one observation, this is a reference and seed dataset, not a large training set. Most reasoning rows are retrieval rather than multi-step computation. 10pages, 2 tables

Visit

doi.org

Tasks

language modeling

Tags

Artificial Intelligence (cs.AI)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Rethinking the Multilingual Reasoning Gap with Layer SwapA Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource LanguagesKinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generationkkrish-tech/low-resource-language-reasoningTIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language ModelsEarly-Layer LoRA Fine-Tuning for Cross-Domain Alignment in Lugha-Llama for Low-Resource Bantu Languages

Rethinking the Multilingual Reasoning Gap with Layer Swap

Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, ev

A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages

Unsupervised part of speech (POS) tagging is often framed as a clustering problem, but practical tag

KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation

The recent mainstream adoption of large language model (LLM) technology is enabling novel applicatio

kkrish-tech/low-resource-language-reasoning

Evaluation framework for measuring LLM reasoning and translation performance across low-resource lan

TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models

To address the severe data scarcity in Tibetan, a low-resource language spoken by over six million p

Early-Layer LoRA Fine-Tuning for Cross-Domain Alignment in Lugha-Llama for Low-Resource Bantu Languages

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their performance in low