Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Low-Resource Corpus Indonesian Local Language

Domain:

natural language processing

Record type:

dataset
Creator:
AhdWibPra
Editor:
Uni
Publisher:
Men
Host:avatar
This study departs from the hypothesis that combining Neural Machine Translation (NMT) with the stemming algorithms ECS, Porter, and Porter Hybrid (ECS) can improve root-word extraction accuracy for Indonesian regional languages. In particular, the hybrid model which integrates ECS’s morphological sensitivity with Porter’s efficiency is expected to deliver the best performance in preserving semantic equivalence across word forms. The dataset comprises nine text files representing three regional languages (Javanese, Minangkabau, and Sundanese) and three stemming approaches (ECS, Porter, and Porter Hybrid-ECS). Each file contains pairs of source sentences and outputs produced through the NMT → Stemming → NMT pipeline: regional-language text is translated into Indonesian, stemmed, and then projected back into the original regional language. The parallel corpora were aligned at the lexical level and curated to balance token counts, structural variation, and morphological diversity. Experimental results reveal several prominent findings. ECS excels for languages with complex affixation because it captures local morphological patterns; Porter is computationally lighter but tends to under-stem prefixes/infices typical of regional languages; and Porter Hybrid (ECS) consistently performs best across all three languages, yielding an average accuracy gain of approximately 4–7% over single models. Evaluation uses precision–recall (assessing correctness and coverage of recovered lemmas), cosine similarity (semantic proximity to reference forms), and BLEU score (n-gram similarity to reference). Interpretively, the data indicate that a hybrid approach blending linguistic insight (ECS) with a general algorithm (Porter) within an NMT-based pipeline produces a more universal and adaptive stemming system for Indonesian regional languages. These findings are relevant for developing cross-regional NLP applications (machine translation, sentiment analysis, lexical normalization), training models for low-resource languages, and improving text preprocessing in translation systems to achieve more accurate morphology-level handling.

Visit

doi.orgdata.mendeley.com

Tasks

machine translationtext normalization

Tags

Natural Language Processing

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Low-Resource Indonesian Health ConversationsConstructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local LanguagesNusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language ModelsA Very Low Resource Language Speech Corpus for Computational Language Documentation ExperimentsTowards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource LanguageRIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

Low-Resource Indonesian Health Conversations

The dataset consists of 5,000 question–answer (QA) pairs collected and annotated for low-resource In

Constructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local Languages

In Indonesia, local languages play an integral role in the culture. However, the available Indonesia

NusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language Models

Large Language Models (LLMs) have demonstrated exceptional promise in translation tasks for high-res

A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments

Most speech and language technologies are trained with massive amounts of speech and text informatio

Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language

The increase in technological adoption worldwide comes with demands for novel tools to be used by th

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

This dataset consists of a curated collection of high-fidelity, field-recorded audio samples develop