Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AM2iCo: Evaluating Word Meaning in Context across Low-Resource Languages with Adversarial Examples

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
LiuPonMcCVul
Hôte:avatar
Capturing word meaning in context and distinguishing between correspondences and variations across languages is key to building successful multilingual and cross-lingual text representation models. However, existing multilingual evaluation datasets that evaluate lexical semantics "in-context" have various limitations. In particular, 1) their language coverage is restricted to high-resource languages and skewed in favor of only a few language families and areas, 2) a design that makes the task solvable via superficial cues, which results in artificially inflated (and sometimes super-human) performances of pretrained encoders, on many target languages, which limits their usefulness for model probing and diagnostics, and 3) little support for cross-lingual evaluation. In order to address these gaps, we present AM2iCo (Adversarial and Multilingual Meaning in Context), a wide-coverage cross-lingual and multilingual evaluation set; it aims to faithfully assess the ability of state-of-the-art (SotA) representation models to understand the identity of word meaning in cross-lingual contexts for 14 language pairs. We conduct a series of experiments in a wide range of setups and demonstrate the challenging nature of AM2iCo. The results reveal that current SotA pretrained encoders substantially lag behind human performance, and the largest gaps are observed for low-resource languages and languages dissimilar to English. EMNLP 2021 long paper

Visit

arxiv.org

Tasks

embeddings

Tags

Computation and Language

Similaires

Context Models for OOV Word Translation in Low-Resource LanguagesRobustness of Cross-Lingual Models to Adversarial Examples in Low-Resource Languages After Intermediate-Task TrainingDependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource LanguagesAdversarial Text-to-Speech for low-resource languagesAdversarial Training for Robust Euphemism Detection in Low-Resource LanguagesMachine Translation in Low-Resource Languages by an Adversarial Neural Network

Context Models for OOV Word Translation in Low-Resource Languages

Out-of-vocabulary word translation is a major problem for the translation of low-resource languages

Robustness of Cross-Lingual Models to Adversarial Examples in Low-Resource Languages After Intermediate-Task Training

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages

Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, ye

Adversarial Text-to-Speech for low-resource languages

Improving the adversarial TTS models for low-resource languages by utilizing the high-frequency similarities between the different languages.

Adversarial Training for Robust Euphemism Detection in Low-Resource Languages

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

Machine Translation in Low-Resource Languages by an Adversarial Neural Network

Existing Sequence-to-Sequence (Seq2Seq) Neural Machine Translation (NMT) shows strong capability wit