Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Cross-Lingual Information Access in the LLM Era: Architectures, Alignment Strategies, and Open Challenges for Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
SidGanSunShr
Éditeur:
Spr
Hôte:
Abstract The way the system is designed to help people find information in languages has changed a lot. We used to rely on translation systems. Now we use neural networks and big language models that can understand many languages. This paper looks at how we got to this point and how these new systems work. We looked at some ideas from the past like the theory of K-representations and SK-languages, which were developed by Fomichov. We also looked at how people used to search for information using ontologies and how they expanded their searches using models. In addition, we considered how people queried information in languages on the web and how they used multilingual frameworks. We used three datasets to test our ideas: MIRACL, XLM RoBERTa benchmarks and NoMIRACL. What we found was interesting. We saw that there is a difference in how well these systems work for languages that have a lot of resources versus those that do not. We also found that the size of the Wikipedia corpus for a language does not necessarily determine how well the system works. Furthermore, we saw that the rates at which these systems make mistakes vary greatly across languages from 0.21 for English to 0.89 for Yoruba. This means that these systems work differently for different languages. Our paper suggests that we need to consider three things when we design these systems: how transparent they are, how well they understand the meaning of words across languages and how fair they are, to all languages. We think that these three things are equally important and that they can help us make systems that work better for everyone. This is an extension of the idea of a language that Fomichov proposed but now we are applying it to the new systems that use big language models.

Visit

doi.org

Tasks

information retrieval

Languages

Yoruba

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Multimodal Alignment for Robust Cross-Lingual NER in Low-Resource LanguagesArtificial Code-Switching for Cross-Lingual Embedding Alignment in Low-Resource LanguagesEnhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word AlignmentCross-lingual NER Generalization via Embedding Alignment in Low-Resource LanguagesOptimal Transport-Based Feature Alignment for Cross-Lingual Retrieval in Low-Resource LanguagesSynthetic Noise Impact on Cross-Lingual NER Alignment in Low-Resource Languages

Multimodal Alignment for Robust Cross-Lingual NER in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Artificial Code-Switching for Cross-Lingual Embedding Alignment in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

The field of cross-lingual sentence embeddings has recently experienced significant advancements, bu

Cross-lingual NER Generalization via Embedding Alignment in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Optimal Transport-Based Feature Alignment for Cross-Lingual Retrieval in Low-Resource Languages

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Synthetic Noise Impact on Cross-Lingual NER Alignment in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident