Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages

Domaine:

natural language processing

Type de record:

paper
We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages. We then compare the performance of OSCAR-based and Wikipedia-based ELMo embeddings for these languages on the part-of-speech tagging and parsing tasks. We show that, despite the noise in the Common-Crawl-based OSCAR data, embeddings trained on OSCAR perform much better than monolingual embeddings trained on Wikipedia. They actually equal or improve the current state of the art in tagging and parsing for all five languages. In particular, they also improve over multilingual Wikipedia-based contextual embeddings (multilingual BERT), which almost always constitutes the previous state of the art, thereby showing that the benefit of a larger, more diverse corpus surpasses the cross-lingual benefit of multilingual embedding architectures.

Visit

arxiv.orgaclanthology.org

Connected records

datasetpaper

Tasks

language modelingembeddings

Languages

AfrikaansAmharicArabic, Egyptian SpokenSomaliSwahiliYoruba

Tags

oscar

Licenses

These data are released under this licensing scheme We do not own any of the text from which these data has been extracted. We license the actual packaging of these data under the Creative Commons CC0 license ("no rights reserved") http://creativecommons.org/publicdomain/zero/1.0/ To the extent possible under law, Inria has waived all copyright and related or neighboring rights to OSCAR This work is published from: France. Should you consider that our data contains material that is owned by you and should therefore not be reproduced here, please: * Clearly identify yourself, with detailed contact data such as an address, telephone number or email address at which you can be contacted. * Clearly identify the copyrighted work claimed to be infringed. * Clearly identify the material that is claimed to be infringing and information reasonably sufficient to allow us to locate the material. We will comply to legitimate requests by removing the affected sources from the next release of the corpus.

Similaires

Multilingual acoustic word embeddings for zero-resource languagesMorphological Segmentation to Improve Crosslingual Word Embeddings for Low Resource LanguagesMultilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related LanguagesPrompt and circumstance: A word-by-word LLM prompting approach to interlinear glossing for low-resource languagesDetecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like BiasesTowards Better UD Parsing: Deep Contextualized Word Embeddings, Ensemble, and Treebank Concatenation

Multilingual acoustic word embeddings for zero-resource languages

This research addresses the challenge of developing speech applications for zero-resource languages

Morphological Segmentation to Improve Crosslingual Word Embeddings for Low Resource Languages

Crosslingual word embeddings developed from multiple parallel corpora help in understanding the rela

Multilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related Languages

Very low-resource languages, having only a few million tokens worth of data, are not well-supported

Prompt and circumstance: A word-by-word LLM prompting approach to interlinear glossing for low-resource languages

Partly automated creation of interlinear glossed text (IGT) has the potential to assist in linguisti

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

With the starting point that implicit human biases are reflected in the statistical regularities of

Towards Better UD Parsing: Deep Contextualized Word Embeddings, Ensemble, and Treebank Concatenation

This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual P