Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

Domain:

natural language processing

Record type:

paperdataset
Creator:
ChhPo,RosCho
Host:avatar
In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Augmented Generation (RAG) framework applied to Khmer agricultural documents. The document chunks are encoded using the BGE-M3 multilingual embedding model and retrieved using the FAISS library. Performance is evaluated using four metrics: Average Retrieval Score (L2 distance), Answer Relevance, Khmer Coverage, and Khmer Intersection over Union, all measured against ground-truth question-answer pairs. For evaluation, we perform 5-fold cross-validation over 18 question-answer pairs. We observe the best performance for the character-based Recursive chunking method with a chunk size of 300 characters, achieving the lowest L2 distance (0.4295 +- 0.0461), highest Answer Relevance (0.8663 +- 0.0199), and highest Khmer IoU (0.6441 +- 0.0347). A paired t-test shows a statistically significant improvement over the Sentence-Based chunking method in L2 distance (p = 0.0121). These results highlight the importance of segmentation granularity and structural preservation for optimizing dense retrieval in morphologically complex, low-resource languages such as Khmer. 11 pages, 1 figure

Visit

arxiv.org

Tasks

embeddingsinformation retrieval

Tags

Computation and LanguageH.3.3; I.2.7

Similar

Optimal Transport Distillation for Low-Resource Language Embedding AlignmentText Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: TigrinyaEvaluation of feature-embedding methods for word spotting in historical arabic documentsLGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language AdaptationEffective Transfer Learning for Low-Resource Natural Language UnderstandingStrategies for improving low resource speech to text translation relying on pre-trained ASR models

Optimal Transport Distillation for Low-Resource Language Embedding Alignment

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Text Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: Tigrinya

This article studies convolutional neural networks for Tigrinya (also referred to as Tigrigna), whic

Evaluation of feature-embedding methods for word spotting in historical arabic documents

International audience Retrieving and indexing historical Arabic documents remain a v

LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation

Adapting pretrained language models to low-resource, morphologically rich languages remains a signif

Effective Transfer Learning for Low-Resource Natural Language Understanding

Natural language understanding (NLU) is the task of semantic decoding of human languages by machines

Strategies for improving low resource speech to text translation relying on pre-trained ASR models

This paper presents techniques and findings for improving the performance of low-resource speech to