Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Cross-Lingual NER for Financial Transaction Data in Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
KumLiuBou
Hôte:avatar
We propose an efficient modeling framework for cross-lingual named entity recognition in semi-structured text data. Our approach relies on both knowledge distillation and consistency training. The modeling framework leverages knowledge from a large language model (XLMRoBERTa) pre-trained on the source language, with a student-teacher relationship (knowledge distillation). The student model incorporates unsupervised consistency training (with KL divergence loss) on the low-resource target language. We employ two independent datasets of SMSs in English and Arabic, each carrying semi-structured banking transaction information, and focus on exhibiting the transfer of knowledge from English to Arabic. With access to only 30 labeled samples, our model can generalize the recognition of merchants, amounts, and other fields from English to Arabic. We show that our modeling approach, while efficient, performs best overall when compared to state-of-the-art approaches like DistilBERT pre-trained on the target language or a supervised model directly trained on labeled data in the target language. Our experiments show that it is enough to learn to recognize entities in English to reach reasonable performance in a low-resource language in the presence of a few labeled samples of semi-structured data. The proposed framework has implications for developing multi-lingual applications, especially in geographies where digital endeavors rely on both English and one or more low-resource language(s), sometimes mixed with English or employed singly. 5 pages, 3 figures. Presented at the SIGIR 2023 Workshop on Knowledge Discovery from Unstructured Data in Financial Services (KDF)

Visit

arxiv.org

Tasks

information extractionnamed entity recognitiontransfer learning

Tags

Computation and Language

Similaires

Cross-lingual NER Model Robustness in Low-Resource LanguagesInference Efficiency Tradeoff in Cross-Lingual NER for Low-Resource LanguagesMultimodal Alignment for Robust Cross-Lingual NER in Low-Resource LanguagesDomain Adaptation Effects in Cross-Lingual NER for Low-Resource LanguagesMultimodal Embedding Integration for Cross-Lingual NER in Low-Resource LanguagesScaling Cross-Lingual NER Performance with Unlabeled Target Data in Low-Resource Languages

Cross-lingual NER Model Robustness in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages

Inference Efficiency Tradeoff in Cross-Lingual NER for Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Multimodal Alignment for Robust Cross-Lingual NER in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Domain Adaptation Effects in Cross-Lingual NER for Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Multimodal Embedding Integration for Cross-Lingual NER in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Scaling Cross-Lingual NER Performance with Unlabeled Target Data in Low-Resource Languages

To better tackle the named entity recognition (NER) problem on languages with little/no labeled data