Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Enhancing Arabic Extractive Summarization with TF-IDF-Weighted AraBERT Sentence Embeddings and Semantic Clustering

Domain:

natural language processing

Record type:

paper
Creator:
WadSurFah
Publisher:
STM
Host:
The increasing amount of textual content across digital platforms, including social media, news and education, has made it difficult for users to extract useful information efficiently. Therefore, Automatic Text Summarization (ATS) becomes an essential tool for distilling large amount of information while maintaining the core idea. Progress in Arabic ATS remains limited due to the scarcity of annotated datasets, the lack of Arabic-specific NLP tools and the high computational cost of LLM. Additionally, traditional methods often fail to capture sentence-level semantics, limiting summary quality. To address this, we propose a scalable, unsupervised framework that uses TF-IDF-weighted AraBERT embeddings to generate rich sentence representations. To further capture document structure, sentences are grouped using k-means clustering. From each cluster, we identify the most representative sentences using centroid similarity and apply Maximal Marginal Relevance (MMR) as a post-processing redundancy to eliminate sentences that are too similar. Experimental evaluation on the EASC dataset demonstrates that our weighted AraBERT model outperforms traditional embedding techniques such as FastText and Unweighted AraBERT, achieving significant improvements across multiple ROUGE metrics.

Visit

doi.org

Tasks

summarizationnatural language generation

Licenses

https://creativecommons.org/licenses/by-sa/4.0

Similar

Uzbek text summarization based on TF-IDFArabic Extractive Summarization Using Pre-Trained ModelsSentence-Level Enhanced Graph-BERT (SEG-BERT) for Extractive Text SummarizationSomali Extractive Text SummarizationExtractive Text Summarization for WolaitaPredictive modelling for fake news detection using TF-IDF and count vectorizers

Uzbek text summarization based on TF-IDF

The volume of information is increasing at an incredible rate with the rapid development of the Inte

Arabic Extractive Summarization Using Pre-Trained Models

Automatic Text Summarization (ATS) is a crucial area of study in Natural Language Processing (NLP) d

Sentence-Level Enhanced Graph-BERT (SEG-BERT) for Extractive Text Summarization

International audience Background: The exponential growth of digital text across news

Somali Extractive Text Summarization

Extractive Text Summarization for Wolaita

Text summarization is the mechanism of summarizing a huge document comprising vast amount of informa

Predictive modelling for fake news detection using TF-IDF and count vectorizers