Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

KrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural Advisory

Domain:

agriculturenatural language processing

Record type:

datasetpaper
Creator:
RezSha
Publisher:
arXiv
Host:avatar
We present KrishokChat, the first citation-grounded Bengali agricultural instruction-tuning dataset for crop advisory in low-resource settings. We establish a foundation of 290 hierarchical Knowledge Nodes, extracting disease symptoms, management practices, chemical dosages, and verbatim citations from 129 domain-filtered agricultural manuals. Every training instance inherits a verified citation header, guaranteeing 100% citation provenance. Using a Partitioned Seed Generation Matrix, these nodes are expanded into 139,200 supervised fine-tuning pairs, and augmented with 5,300 chemical safety and 1,000 adversarial safety instances, yielding 145,500 QA pairs across 18 crop categories. To evaluate real-world performance, we introduce the Farmer Benchmark, comprising 1,001 authentic farmer queries curated from field surveys and digital portals. Empirical evaluation on Gemma-4-E2B reveals that while fine-tuning on KrishokChat vastly improves structured formatting, standalone models still struggle with exact chemical dosage generalization. This highlights the dataset's true value as a verified knowledge base for retrieval-augmented generation (RAG) rather than mere parametric memorization. All data, code, and benchmarks are released under CC-BY-4.0.

Visit

doi.orgarxiv.org

Tasks

question answering

Tags

Machine Learning (cs.LG)FOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similar

Uganda Agricultural Advisory Dataset (UG-AgriAdvisory)Cost-Efficient Cross-Lingual Retrieval-Augmented Generation for Low-Resource Languages: A Case Study in Bengali Agricultural AdvisoryBanHealthAIRespRelv3: A Benchmark Dataset for Evaluating Bengali AI-Generated Health Responses Based on User-Specific PromptsTukaBench: A Culturally Grounded Jailbreak Benchmark for African LanguagesBanglaFakeNews: A Curated Dataset for Bengali Fake News DetectionPatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

Uganda Agricultural Advisory Dataset (UG-AgriAdvisory)

Built with Adaptive Data by Adaption | Crane AI Labs Submitted to the Uncharted Data Challenge 2026

Cost-Efficient Cross-Lingual Retrieval-Augmented Generation for Low-Resource Languages: A Case Study in Bengali Agricultural Advisory

Access to reliable agricultural advisory remains limited in many developing regions due to a persist

BanHealthAIRespRelv3: A Benchmark Dataset for Evaluating Bengali AI-Generated Health Responses Based on User-Specific Prompts

BanHealthAIRespRelv3 serves as a benchmark dataset for evaluating the relevance and safety of AI-gen

TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resou

BanglaFakeNews: A Curated Dataset for Bengali Fake News Detection

BanglaFakeNews is a large-scale, curated dataset developed for fake news detection in the Bengali la

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language underst