Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Loanword Identification in Low-Resource Languages with Minimal Supervision

Domain:

natural language processing

Record type:

paper
Creator:
CheLei XieYan
Publisher:
Ass
Host:
Bilingual resources play a very important role in many natural language processing tasks, especially the tasks in cross-lingual scenarios. However, it is expensive and time consuming to build such resources. Lexical borrowing happens in almost every language. This inspires us to detect these loanwords effectively, and to use the “loanword (in receipt language)”-“donor word (in donor language)” to extend the bilingual resource for NLP tasks in low-resource languages. In this article, we propose a novel method to identify loanwords in Uyghur. The most important advantage of this method is that the model only relies on large amount of monolingual corpora and only a small scale of annotated data. Our loanword identification model includes two parts: loanword candidate generation and loanword prediction. In the first part, we use two large-scale monolingual corpora and a small bilingual dictionary to train a cross-lingual embedding model. Since semantic unrelated words often cannot be treated as loanword pairs, a loanword candidate list will be generated according to this model and a word list in Uyghur. In the second part, we predict from the preceding candidates based on a log-linear model that integrates several features such as pronunciation similarity, part-of-speech tags, and hybrid language modeling. To evaluate the effectiveness of our proposed method, we conduct two types of experiments: loanword identification and OOV translation. Experimental results showed that (1) our proposed method achieved significant F1 improvements compared to other models in all four loanword identification tasks in Uyghur, and (2) after extending the existing translation models with loanword identification results, OOV rates in several language pairs reduced significantly and the translation performance improved.

Visit

doi.org

Licenses

https://www.acm.org/publications/policies/copyright_policy#Background

Similar

Extracting Parallel Sentences from Low-Resource Language Pairs with Minimal SupervisionUnlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training DataGlotLID: Language Identification for Low-Resource LanguagesHomophobia and transphobia span identification in low-resource languagesSoumyaTeotia/Cross-Domain-Multi-Intent-Identification-in-Low-Resource-Languagesdev52003/Biased-News-Identification-For-Low-Resource-Languages-Malayalam

Extracting Parallel Sentences from Low-Resource Language Pairs with Minimal Supervision

Abstract At present, machine translation in the market depends on parallel sentence

Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data

Recent advances in LLMs have enhanced AI capabilities, but also increased the risk posed by maliciou

GlotLID: Language Identification for Low-Resource Languages

International audience Several recent papers have published good solutions for langua

Homophobia and transphobia span identification in low-resource languages

Online platforms have become prevalent because they promote free speech and group discussions. Howev

SoumyaTeotia/Cross-Domain-Multi-Intent-Identification-in-Low-Resource-Languages

Addressing the challenges of cross-domain and multi-intent classification in low-resource languages,

dev52003/Biased-News-Identification-For-Low-Resource-Languages-Malayalam

# 📰 Biased News Identification in Malayalam Media This repository contains the research work and to