Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A neural approach for inducing multilingual resources and natural language processing tools for low-resource languages

Domain:

natural language processing

Record type:

paper
Creator:
O. N. L.
Publisher:
Cam
Host:
Abstract This work focuses on the rapid development of linguistic annotation tools for low-resource languages (languages that have no labeled training data). We experiment with several cross-lingual annotation projection methods using recurrent neural networks (RNN) models. The distinctive feature of our approach is that our multilingual word representation requires only a parallel corpus between source and target languages. More precisely, our approach has the following characteristics: (a) it does not use word alignment information, (b) it does not assume any knowledge about target languages (one requirement is that the two languages (source and target) are not too syntactically divergent), which makes it applicable to a wide range of low-resource languages, (c) it provides authentic multilingual taggers (one tagger for N languages). We investigate both uni and bidirectional RNN models and propose a method to include external information (for instance, low-level information from part-of-speech tags) in the RNN to train higher level taggers (for instance, Super Sense taggers). We demonstrate the validity and genericity of our model by using parallel corpora (obtained by manual or automatic translation). Our experiments are conducted to induce cross-lingual part-of-speech and Super Sense taggers. We also use our approach in a weakly supervised context, and it shows an excellent potential for very low-resource settings (less than 1k training utterances).

Visit

doi.org

Tasks

part of speech taggingtransfer learning

Licenses

https://www.cambridge.org/core/terms

Similar

Natural Language Processing (NLP) tools - multilingual and low-resource languagesCharacter-level and syntax-level models for low-resource and multilingual natural language processingEfficient and human-inspired natural language processing methods for multilingual and low-resource settingsMultilingual Neural Machine Translation for Low Resource LanguagesNatural Language Processing in Low-Resource Languages: Progress and ProspectsBuilding natural language processing tools for Runyakitara

Natural Language Processing (NLP) tools - multilingual and low-resource languages

Character-level and syntax-level models for low-resource and multilingual natural language processing

There are more than 7000 languages in the world, but only a small portion of them benefit from Natur

Efficient and human-inspired natural language processing methods for multilingual and low-resource settings

The rapid advancement of large language models (LLMs) has revolutionized natural language processing

Multilingual Neural Machine Translation for Low Resource Languages

Neural Machine Translation (NMT) has been shown to be more effective in translation tasks compared t

Natural Language Processing in Low-Resource Languages: Progress and Prospects

Low-resource languageslanguages with limited annotated corpora, lexicons, and digital resourcespose

Building natural language processing tools for Runyakitara

Abstract This paper describes an endeavour to build natural language processing (NL