International audience
Cross-lingual summarization aims to condense written or spoken content in one language into a coherent and concise summary in another language. This task requires understanding the source language's nuances, filtering for importance, and reconstructing meaning in a different grammar and vocabulary; these operations are particularly difficult to perform when the available data are scarce to train models. In this paper, we propose a novel framework that uses multilingual sentence and utterance embeddings to process both speech and text inputs under strict data constraints. Based on the CrossSum dataset and a new cross-lingual speech evaluation dataset we collected for three low-resource languages, our experimental results highlight the strong potential of our approach for cross-lingual summarization, particularly in low-resource spoken language scenarios.