Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Cross-Lingual Dialogue Dataset Creation via Outline-Based Generation

Domain:

natural language processing

Record type:

paperdataset
Creator:
MajRazPonVul
Host:avatar
Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (communities of) speakers. Nevertheless, the potential of this technology is not fully realised, as current datasets for multilingual ToD - both for modular and end-to-end modelling - suffer from severe limitations. 1) When created from scratch, they are usually small in scale and fail to cover many possible dialogue flows. 2) Translation-based ToD datasets might lack naturalness and cultural specificity in the target language. In this work, to tackle these limitations we propose a novel outline-based annotation process for multilingual ToD datasets, where domain-specific abstract schemata of dialogue are mapped into natural language outlines. These in turn guide the target language annotators in writing a dialogue by providing instructions about each turn's intents and slots. Through this process we annotate a new large-scale dataset for training and evaluation of multilingual and cross-lingual ToD systems. Our Cross-lingual Outline-based Dialogue dataset (termed COD) enables natural language understanding, dialogue state tracking, and end-to-end dialogue modelling and evaluation in 4 diverse languages: Arabic, Indonesian, Russian, and Kiswahili. Qualitative and quantitative analyses of COD versus an equivalent translation-based dataset demonstrate improvements in data quality, unlocked by the outline-based approach. Finally, we benchmark a series of state-of-the-art systems for cross-lingual ToD, setting reference scores for future work and demonstrating that COD prevents over-inflated performance, typically met with prior translation-based ToD datasets.

Visit

arxiv.org

Languages

SwahiliSwahili, CoastalSwahili, Congo

Tags

Computation and Language

Similar

AfriWOZ: Corpus for Exploiting Cross-Lingual Transfer for Dialogue Generation in Low-Resource, African LanguagesOverlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related LanguagesXGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and GenerationUnsupervised ASR via Cross-Lingual Pseudo-LabelingCross-lingual Representation Learning via Centroid Intervention FusionMitigating Annotation Projection Noise in Cross-Lingual NER via Source Dataset Scaling for Low-Resource Languages

AfriWOZ: Corpus for Exploiting Cross-Lingual Transfer for Dialogue Generation in Low-Resource, African Languages

Overlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related Languages

Pre-trained multilingual language models such as mBERT and XLM-R have demonstrated great potential f

XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation

In this paper, we introduce XGLUE, a new benchmark dataset that can be used to train large-scale cross-lingual pre-trained models using multilingual and bilingual corpora and evaluate their performance across a diverse set of cross-lingual tasks. Comparing to GLUE(

Unsupervised ASR via Cross-Lingual Pseudo-Labeling

Recent work has shown that it is possible to train an $\textit{unsupervised}$ automatic speech recog

Cross-lingual Representation Learning via Centroid Intervention Fusion

Large language models (LLMs) exhibit uneven multilingual performance, especially when dealing with l

Mitigating Annotation Projection Noise in Cross-Lingual NER via Source Dataset Scaling for Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident