Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Bridging Languages and Modalities: Lightweight Cross-Lingual Text and Speech Summarization for Low-Resource Scenarios

Domain:

natural language processing

Record type:

paper
Creator:
CheMdhEstHue
Editor:
LabLab
Publisher:
CCSD
Host:avatar
International audience Cross-lingual summarization aims to condense written or spoken content in one language into a coherent and concise summary in another language. This task requires understanding the source language's nuances, filtering for importance, and reconstructing meaning in a different grammar and vocabulary; these operations are particularly difficult to perform when the available data are scarce to train models. In this paper, we propose a novel framework that uses multilingual sentence and utterance embeddings to process both speech and text inputs under strict data constraints. Based on the CrossSum dataset and a new cross-lingual speech evaluation dataset we collected for three low-resource languages, our experimental results highlight the strong potential of our approach for cross-lingual summarization, particularly in low-resource spoken language scenarios.

Visit

hal.science

Tasks

summarizationnatural language generation

Tags

low-resource languagesutterance embeddingssentence embeddingscrosslingual summarization[INFO]Computer Science [cs]

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource LanguagesXF2T: Cross-lingual Fact-to-Text Generation for Low-Resource LanguagesText-to-speech system for low-resource language using cross-lingual transfer learning and data augmentationAdversarial Text-to-Speech for low-resource languagesOmnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and SpeechEffectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognition

XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource Languages

Lack of encyclopedic text contributors, especially on Wikipedia, makes automated text generation for

XF2T: Cross-lingual Fact-to-Text Generation for Low-Resource Languages

Multiple business scenarios require an automated generation of descriptive human-readable text from

Text-to-speech system for low-resource language using cross-lingual transfer learning and data augmentation

Abstract Deep learning techniques are currently being applied in automated text-to-speech (TTS) sys

Adversarial Text-to-Speech for low-resource languages

Improving the adversarial TTS models for low-resource languages by utilizing the high-frequency similarities between the different languages.

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech

Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstr

Effectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognition

In the recent years end to end (E2E) automatic speech recognition (ASR) systems have achieved promis