Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Improving Multilingual Semantic Textual Similarity with Shared Sentence Encoder for Low-resource Languages

Domain:

natural language processing

Record type:

papermodel
Creator:
TanCheDo,Min
Host:avatar
Measuring the semantic similarity between two sentences (or Semantic Textual Similarity - STS) is fundamental in many NLP applications. Despite the remarkable results in supervised settings with adequate labeling, little attention has been paid to this task in low-resource languages with insufficient labeling. Existing approaches mostly leverage machine translation techniques to translate sentences into rich-resource language. These approaches either beget language biases, or be impractical in industrial applications where spoken language scenario is more often and rigorous efficiency is required. In this work, we propose a multilingual framework to tackle the STS task in a low-resource language e.g. Spanish, Arabic , Indonesian and Thai, by utilizing the rich annotation data in a rich resource language, e.g. English. Our approach is extended from a basic monolingual STS framework to a shared multilingual encoder pretrained with translation task to incorporate rich-resource language data. By exploiting the nature of a shared multilingual encoder, one sentence can have multiple representations for different target translation language, which are used in an ensemble model to improve similarity evaluation. We demonstrate the superiority of our method over other state of the art approaches on SemEval STS task by its significant improvement on non-MT method, as well as an online industrial product where MT method fails to beat baseline while our approach still has consistently improvements.

Visit

arxiv.org

Tasks

embeddings

Tags

Computation and Language

Similar

A Resource-Light Method for Cross-Lingual Semantic Textual SimilarityPerformance of XLM-R Models on Zero-Shot Cross-Lingual Semantic Textual Similarity in XTREME-R for Low-Resource African LanguagesRobust Multilingual Encoder Training for Low-Resource Language AlignmentLeveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource LanguagesImproving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity AnalysisEnhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

A Resource-Light Method for Cross-Lingual Semantic Textual Similarity

Recognizing semantically similar sentences or paragraphs across languages is beneficial for many tas

Performance of XLM-R Models on Zero-Shot Cross-Lingual Semantic Textual Similarity in XTREME-R for Low-Resource African Languages

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

Robust Multilingual Encoder Training for Low-Resource Language Alignment

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

Leveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource Languages

The importance of qualitative parallel data in machine translation has long been determined but it h

Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis

This paper examines how linguistic similarity affects cross-lingual phonetic representation in speec

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

The field of cross-lingual sentence embeddings has recently experienced significant advancements, bu