Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Towards Globally Inclusive Multilingual Dialogue Systems for Real-World Applications

Domaine:

natural language processing
Créateur:
Hu,
Éditeur:
ApoUniKorVul
Éditeur:
Apo
Hôte:avatar
With the advent of large language models (LLMs), dialogue systems have become the primary interface for accessing advances in natural language processing (NLP); yet existing research remains largely English-centric, text-based, and benchmark-driven, limiting both global inclusivity and real-world applications. To move beyond this narrow focus, the thesis broadens the scope of multilingual dialogue system research through the creation of new datasets and evaluation methods. It introduces Multi3WOZ, a large-scale, multi-parallel dataset for task-oriented dialogue in Arabic, English, French, and Turkish. It also presents HEALTHDIAL, the first large-scale, speech-first dataset for health communication, spanning diverse language varieties across Arabic, Chinese, English, and Spanish. HEALTHDIAL modernises task-oriented dialogue system design by replacing traditional parsing-based approaches with a retrieval-augmented generation pipeline that more effectively leverages the capabilities of LLMs. In addition, the thesis proposes the first framework for the quantitative measurement of cross-lingual disparities, capturing both those arising during system development and those intrinsic to LLMs. The conventional dataset--model--benchmarking pipeline has driven much of the progress in dialogue system research, but it remains insufficient for informing real-world applications. To bridge this gap, the thesis extends the pipeline in both directions. Upstream, it applies systematic review methodology in combination with global health frameworks to identify user needs, and map the state of NLP for public health in Africa. Downstream, it develops and releases open-source toolkits for multilingual data collection, system development, deployment, and human evaluation, thereby lowering barriers to real-world applications. Beyond its technical contributions, this thesis offers a methodological reflection on how NLP, and dialogue system research in particular, can move beyond benchmarks to generate evidence with real-world relevance. It distils three guiding principles for equitable NLP: research should be evidence-based, grounding decisions in systematic evidence; human-centric, ensuring that development and evaluation reflect the needs and values of the communities served; and context-adaptive, responding to the resources and constraints of diverse linguistic and cultural contexts. Together, these principles outline a framework for developing dialogue systems that are both globally inclusive and socially impactful.

Visit

doi.orgwww.repository.cam.ac.uk

Tags

Evaluation MethodologyLarge Language ModelsMultilingual Dialogue SystemsNatural Language ProcessingSpoken Dialogue Systems

Licenses

All rights reservedhttp://purl.org/NET/rdflicense/allrightsreservedopen.accesshttp://purl.org/coar/access_right/c_abf2

Similaires

ML-Embed: Inclusive and Efficient Embeddings for a Multilingual WorldTowards Inclusive Knowledge Organizational Systems for Multilingual Communities and Conceptual Models in Federal Universities Libraries in NigeriaTowards real world medical image analysisF2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual WorldTowards Real-World Streaming Speech Translation for Code-Switched SpeechLEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World

ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World

The development of high-quality text embeddings is increasingly drifting toward an exclusionary futu

Towards Inclusive Knowledge Organizational Systems for Multilingual Communities and Conceptual Models in Federal Universities Libraries in Nigeria

Knowledge organization systems (KOS)—including classification schemes, subject heading lists, thesau

Towards real world medical image analysis

Many foundation models for medical image analysis, such as Segment Anything Model (SAM), have been r

F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World

We present F2LLM-v2, a new family of general-purpose, multilingual embedding models in 8 distinct si

Towards Real-World Streaming Speech Translation for Code-Switched Speech

Code-switching (CS), i.e. mixing different languages in a single sentence, is a common phenomenon in

LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World

This paper presents LEMONADE, a large-scale conflict event dataset comprising 39,786 events across 2