Logo Lanfrica

Cross-Lingual Information Access in the LLM Era: Architectures, Alignment Strategies, and Open Challenges for Low-Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
SidGanSunShr
Publisher:
Spr
Host:
Abstract The way the system is designed to help people find information in languages has changed a lot. We used to rely on translation systems. Now we use neural networks and big language models that can understand many languages. This paper looks at how we got to this point and how these new systems work. We looked at some ideas from the past like the theory of K-representations and SK-languages, which were developed by Fomichov. We also looked at how people used to search for information using ontologies and how they expanded their searches using models. In addition, we considered how people queried information in languages on the web and how they used multilingual frameworks. We used three datasets to test our ideas: MIRACL, XLM RoBERTa benchmarks and NoMIRACL. What we found was interesting. We saw that there is a difference in how well these systems work for languages that have a lot of resources versus those that do not. We also found that the size of the Wikipedia corpus for a language does not necessarily determine how well the system works. Furthermore, we saw that the rates at which these systems make mistakes vary greatly across languages from 0.21 for English to 0.89 for Yoruba. This means that these systems work differently for different languages. Our paper suggests that we need to consider three things when we design these systems: how transparent they are, how well they understand the meaning of words across languages and how fair they are, to all languages. We think that these three things are equally important and that they can help us make systems that work better for everyone. This is an extension of the idea of a language that Fomichov proposed but now we are applying it to the new systems that use big language models.

Languages

Similar