Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Building a Dataset and Exploring Low-Resource Approaches to Natural Language Inference with Myanmar

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Hte
Éditeur:
Mac
Hôte:avatar
Despite dramatic recent progress in NLP, it is still a major challenge to apply Large Language Models (LLM) to low-resource languages. This is most visible in benchmarks such as Cross-Lingual Natural Language Inference (XNLI), a key task that demonstrates cross-lingual capabilities of NLP systems across a set of 15 languages. In this thesis, we extend XNLI task for one additional low-resource language, Myanmar, as a proxy challenge for broader low-resource languages, and make three core contributions. First, we build a dataset called Myanmar XNLI (myXNLI) using community crowd-sourced methods, as an extension to the existing XNLI corpus. This involves a two-stage process of community-based construction followed by expert verification; through an analysis, we demonstrate and quantify the value of the expert verification stage in the context of community-based construction for low-resource languages. We make the myXNLI dataset available to the community for future research. Second, we carry out evaluations of recent multilingual language models on the myXNLI benchmark, as well as explore data-augmentation methods to improve model performance. Our data-augmentation methods improve model accuracy by up to 2 percentage points for Myanmar, while uplifting other languages at the same time. Third, we investigate how well these data-augmentation methods generalise to other low-resource languages in the XNLI dataset.

Visit

doi.orgfigshare.mq.edu.au

Tasks

natural language inferencetransfer learning

Tags

Natural language processing

Licenses

In Copyrighthttp://rightsstatements.org/vocab/InC/1.0/

Similaires

JamPatoisNLI: A Jamaican Patois Natural Language Inference DatasetTIGSEN: Building a Low-Resource Dataset and Benchmarking for Tigrigna Sentiment Analysis with Cross-Lingual Transfer Learning ApproachesLow-Resource African Language Pretraining for Zero-Shot XTREME-R Natural Language Inference AccuracyBuilding a Dataset for Misinformation Detection in the Low-Resource LanguageFrom Extractive to Transformer-Based Summarization: A Review of Low-Resource Language Approachesmagnusklundgren/Low-resource-language-dataset

JamPatoisNLI: A Jamaican Patois Natural Language Inference Dataset

JamPatoisNLI provides the first dataset for natural language inference in a creole language, Jamaica

TIGSEN: Building a Low-Resource Dataset and Benchmarking for Tigrigna Sentiment Analysis with Cross-Lingual Transfer Learning Approaches

Low-Resource African Language Pretraining for Zero-Shot XTREME-R Natural Language Inference Accuracy

Multilingual Pretrained Language Models (MPLMs) have shown their strong multilinguality in recent em

Building a Dataset for Misinformation Detection in the Low-Resource Language

From Extractive to Transformer-Based Summarization: A Review of Low-Resource Language Approaches

The digital world is expanding at an unprecedented pace, with recent estimates suggesting that over

magnusklundgren/Low-resource-language-dataset

Repository for evaluating a dataset for low reource languages, based on Amazon's Mintaka dataset #