Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods

Domain:

natural language processing

Record type:

paperdataset
Creator:
GowKarSheDar
Host:avatar
This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pretraining to expand language technologies to understudied and typologically diverse languages, its effectiveness is dependent on quality and suitable benchmarks. We release new sense-annotated datasets of sentences containing polysemous words, spanning ten low-resource languages across diverse language families and scripts. To facilitate dataset creation, the paper presents a demonstrably beneficial semi-automatic annotation method. The utility of the datasets is demonstrated through Word-in-Context (WiC) formatted experiments that evaluate transfer on these low-resource languages. Results highlight the importance of targeted dataset creation and evaluation for effective polysemy disambiguation in low-resource settings and transfer studies. The released datasets and code aim to support further research into fair, robust, and truly multilingual NLP. 8 pages, 22 figures, published at SIGTYP 2025 workshop in ACL

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similar

SenWiCh: Sense-Annotated Sentences for WSD and WiC in Low-Resource LanguagesImproving Resource Creation for Low-Resource Languages using NLP Methods

SenWiCh: Sense-Annotated Sentences for WSD and WiC in Low-Resource Languages

SenWiCh is a multilingual dataset of sense-annotated sentences designed to support research in Word

Improving Resource Creation for Low-Resource Languages using NLP Methods

To digitize existing high-quality text belonging to a certain low-resource language, we are often fa