Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Abstractive Summarization for Urdu Video Description Generation

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
AliFaiMuhAsi
Éditeur:
Ass
Hôte:
Automatic summarization condenses content while retaining key ideas and details. Urdu, with over 230 million speakers globally, is one of the most widely spoken languages. The rise of Urdu content on social media platforms has driven the need for tools that enhance accessibility and engagement. The growing popularity of social media has increased the number of Urdu instructional videos. Well-written video descriptions can boost viewer engagement and improve search engine optimization; however, many lack these. Therefore, an automatic description generation system for Urdu videos is needed, which can be achieved by abstractive summarization of video transcripts. However, such public datasets are not available in Urdu. To address this problem, we investigate the usability of high-resource language datasets for Urdu abstractive text summarization. We created the first Urdu video transcription dataset Urdu How2 and evaluated its quality using intrinsic evaluation. We leverage transfer learning, a technique where knowledge from pretrained models (like mT5) is adapted to new tasks, to develop the uT5 model for generating Urdu text summaries. We further trained the model to improve its Urdu text generation capability. The machine-generated summaries are evaluated using ROUGE scores, human evaluation scores, and adversarial evaluation, providing a reliable assessment of the quality of generated descriptions and the robustness of the model against noisy text data. The human evaluation shows the proposed method generates accurate and coherent summaries compared to the translated ground truth. To the best of our knowledge, this is the first attempt to utilize a cross-lingual dataset for Urdu abstractive text summarization and video description generation. This research enhances Urdu content accessibility and lays the groundwork for advancing multilingual content generation and multimodal analysis in other low-resource languages.

Visit

doi.org

Tasks

natural language generationsummarization

Similaires

Amharic Abstractive Text SummarizationEnhanced model for abstractive Arabic text summarization using natural language generation and named entity recognitionUnsupervised Abstractive Summarization of Bengali Text DocumentsAbstractive Text Summarization for Contemporary Sanskrit Prose: Issues and ChallengesXL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 LanguagesLEMMA-ROUGE: An Evaluation Metric for Arabic Abstractive Text Summarization

Amharic Abstractive Text Summarization

Text Summarization is the task of condensing long text into just a handful of sentences. Many approa

Enhanced model for abstractive Arabic text summarization using natural language generation and named entity recognition

Abstract With the rise of Arabic digital content, effective summarization methods are es

Unsupervised Abstractive Summarization of Bengali Text Documents

Abstractive summarization systems generally rely on large collections of document-summary pairs. How

Abstractive Text Summarization for Contemporary Sanskrit Prose: Issues and Challenges

This thesis presents Abstractive Text Summarization models for contemporary Sanskrit prose. The firs

XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages

Contemporary works on abstractive text summarization have focused primarily on highresource languages like English, mostly due to the limited availability of datasets for low/midresource ones. In this work, we present XLSum, a comprehensive and diverse dataset comp

LEMMA-ROUGE: An Evaluation Metric for Arabic Abstractive Text Summarization

High morphological languages are characterized by complex inflections and derivations, which can pre