Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Building a Digital Lexical Resource for Banyumasan Javanese: A Low-Resource Language Approach

Domaine:

natural language processing

Type de record:

datasetsoftware
Créateur:
NisAngVig
Éditeur:
Uni
Hôte:
Banyumasan Javanese, widely recognized through the Ngapak dialect, remains culturally significant but is still underrepresented in reusable computational resources. This study develops a digital Banyumasan-Indonesian lexical corpus and frames it as a reusable research artifact rather than a static appendix. The corpus was constructed from a Banyumasan-Indonesian dictionary, normalized into a structured bilingual dataset, and packaged as an installable Python resource so that it can be used directly in computational experiments. The implemented resource supports dataset loading, Banyumasan lookup, Indonesian lookup, simple translation, structured translation analysis, batch translation, and corpus statistics. The resulting corpus contains 2,000 lexical pairs, 1,996 unique Banyumasan forms, 1,444 unique Indonesian equivalents, and 4 duplicated Banyumasan headwords that preserve lexical ambiguity from the source material. To demonstrate practical utility, the study includes a 100-sentence implementation example in which Banyumasan text is translated with the published banyumasan-corpus package and evaluated against Indonesian ground truth using the Indonesian-focused embedding model LazarusNLP/all-indo-e5-small-v4. The average semantic similarity rises from 0.4833 for direct Banyumasan-versus-ground-truth comparison to 0.6427 after translation, producing an absolute gain of 0.1594 and a relative improvement of approximately 33.0% over the baseline. These findings indicate that a structured lexical corpus, when distributed in a directly reusable computational form, can strengthen both resource accessibility and small-scale downstream experimentation for a low-resource regional language.

Visit

doi.org

Tasks

machine translation

Licenses

https://creativecommons.org/licenses/by-nc-sa/4.0

Similaires

Building a Rich Lexical Resource for Standard ArabicBuilding a Dataset for Misinformation Detection in the Low-Resource LanguageBuilding an asr system for a low-resource language through the adaptation of a high-resource language asr system: preliminary resultsA Proposed Approach for Extracting Semantic and Lexical Relations for Low-Resource Languages: A Case Study of DarijaTowards Digital Preservation of Efik: TTS for a Low-Resource African LanguageBuilding ‌a ‌Transformer-Based ‌Neural Machine Translation System for English–Kibajuni Translation: A Low-Resource Deep Learning Approach for Indigenous Language Preservation

Building a Rich Lexical Resource for Standard Arabic

Language ambiguity is an inherent characteristic of natural languages. It refers to the phenomenon w

Building a Dataset for Misinformation Detection in the Low-Resource Language

Building an asr system for a low-resource language through the adaptation of a high-resource language asr system: preliminary results

International audience

A Proposed Approach for Extracting Semantic and Lexical Relations for Low-Resource Languages: A Case Study of Darija

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native spe

Building ‌a ‌Transformer-Based ‌Neural Machine Translation System for English–Kibajuni Translation: A Low-Resource Deep Learning Approach for Indigenous Language Preservation

Recent progress in artificial intelligence has pushed machine translation to high levels of accuracy