Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Enhancing Arabic Extractive Summarization with TF-IDF-Weighted AraBERT Sentence Embeddings and Semantic Clustering

Domaine:

natural language processing

Type de record:

paper
Créateur:
WadSurFah
Éditeur:
STM
Hôte:
The increasing amount of textual content across digital platforms, including social media, news and education, has made it difficult for users to extract useful information efficiently. Therefore, Automatic Text Summarization (ATS) becomes an essential tool for distilling large amount of information while maintaining the core idea. Progress in Arabic ATS remains limited due to the scarcity of annotated datasets, the lack of Arabic-specific NLP tools and the high computational cost of LLM. Additionally, traditional methods often fail to capture sentence-level semantics, limiting summary quality. To address this, we propose a scalable, unsupervised framework that uses TF-IDF-weighted AraBERT embeddings to generate rich sentence representations. To further capture document structure, sentences are grouped using k-means clustering. From each cluster, we identify the most representative sentences using centroid similarity and apply Maximal Marginal Relevance (MMR) as a post-processing redundancy to eliminate sentences that are too similar. Experimental evaluation on the EASC dataset demonstrates that our weighted AraBERT model outperforms traditional embedding techniques such as FastText and Unweighted AraBERT, achieving significant improvements across multiple ROUGE metrics.

Visit

doi.org

Tasks

summarizationnatural language generation

Licenses

https://creativecommons.org/licenses/by-sa/4.0

Similaires

Uzbek text summarization based on TF-IDFArabic Extractive Summarization Using Pre-Trained ModelsSentence-Level Enhanced Graph-BERT (SEG-BERT) for Extractive Text SummarizationSomali Extractive Text SummarizationExtractive Text Summarization for WolaitaPredictive modelling for fake news detection using TF-IDF and count vectorizers

Uzbek text summarization based on TF-IDF

The volume of information is increasing at an incredible rate with the rapid development of the Inte

Arabic Extractive Summarization Using Pre-Trained Models

Automatic Text Summarization (ATS) is a crucial area of study in Natural Language Processing (NLP) d

Sentence-Level Enhanced Graph-BERT (SEG-BERT) for Extractive Text Summarization

International audience Background: The exponential growth of digital text across news

Somali Extractive Text Summarization

Extractive Text Summarization for Wolaita

Text summarization is the mechanism of summarizing a huge document comprising vast amount of informa

Predictive modelling for fake news detection using TF-IDF and count vectorizers