Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Morphological Segmentation to Improve Crosslingual Word Embeddings for Low Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
SanDinAsh
Publisher:
Ass
Host:
Crosslingual word embeddings developed from multiple parallel corpora help in understanding the relationships between languages and improving the prediction quality of machine translation. However, in low resource languages with complex and agglutinative morphologies, inducing good-quality crosslingual embeddings becomes challenging due to the problem of complex morphological forms and rare words. This is true even for languages that share common linguistic structure. In our work, we have shown that performing a simple morphological segmentation upon the corpora prior to the generation of crosslingual word embeddings for both roots and suffixes greatly improves the prediction quality and captures semantic similarities more effectively. To exhibit this, we have chosen two related languages: Telugu and Kannada of the Dravidian language family. We have also tested our method upon a widely spoken North Indian language, Hindi, belonging to the Indo-European language family, and have observed encouraging results.

Visit

doi.org

Tasks

embeddings

Licenses

https://www.acm.org/publications/policies/copyright_policy#Background

Similar

Multilingual acoustic word embeddings for zero-resource languagesMultilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related LanguagesA Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesMultilingual Contextual Adapters To Improve Custom Word Recognition In Low-resource LanguagesIsomorphic Cross-lingual Embeddings for Low-Resource LanguagesLearning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel Corpora

Multilingual acoustic word embeddings for zero-resource languages

This research addresses the challenge of developing speech applications for zero-resource languages

Multilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related Languages

Very low-resource languages, having only a few million tokens worth of data, are not well-supported

A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages

We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages. We then compare the performance of OSCAR-based and Wik

Multilingual Contextual Adapters To Improve Custom Word Recognition In Low-resource Languages

Connectionist Temporal Classification (CTC) models are popular for their balance between speed and p

Isomorphic Cross-lingual Embeddings for Low-Resource Languages

Cross-Lingual Word Embeddings (CLWEs) are a key component to transfer linguistic information learnt

Learning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel Corpora

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small