Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Fine‐Tuned Pretrained Transformer for Amharic News Headline Generation

Domaine:

natural language processing

Type de record:

modelpaper
Créateur:
MizMil
Éditeur:
WILEY
Hôte:
ABSTRACT Amharic is one of the under‐resourced languages, making news headline generation particularly challenging due to the scarcity of high‐quality linguistic datasets necessary for training effective natural language processing models. In this study, we fine‐tuned the small check point of the T5v1.1 model (t5‐small) to perform Amharic news headline generation with an Amharic dataset that is comprised of over 70k news articles along with their headline. Fine‐tuning the model involves dataset collection from Amharic news websites, text cleaning, news article size optimization using the TF‐IDF algorithm, and tokenization. In addition, a tokenizer model is developed using the byte pair encoding (BPE) algorithm prior to feeding the dataset for feature extraction and summarization. Metrics including Rouge‐L, BLEU, and Meteor were used to evaluate the performance of the model and a score of 0.5, 0.24, and 0.71, respectively, was achieved on the test partition of the dataset that contains 7230 instances. The results were good relative to result of the t5 model without fine‐tuning, which are 0.1, 0.03, and 0.14, respectively. A postprocessing technique using a rule‐based approach was used for further improving summaries generated by the model. The addition of the postprocessing helped the system to achieve Rouge‐L, BLEU, and Meteor scores of 0.72, 0.52, and 0.81, respectively. The result value is relatively better than the result achieved by the nonfine‐tuned T5v1.1 model and the result of previous studies report on abstractive‐based text summarization for Amharic language, which had a 0.27 Rouge‐L score. This contributes a valuable insight for practical application and further improvement of the model in the future by increasing the article length, using more training data, using machine learning–based adaptive postprocessing techniques, and fine‐tuning other available pretrained models for text summarization.

Visit

doi.org

Tasks

natural language generationsummarization

Languages

Amharic

Licenses

http://creativecommons.org/licenses/by/4.0/

Similaires

Transformer Based Amharic Headline Generation using Sub-word2Vec RepresentationAfriHG: News headline generation for African Languagesaltaseb12/Automatic-Amharic-Text-News-Classification-Using-Fine--Tuned-BERT-ModelsWhen Less Is Enough: Context Selection and Prompting Strategies for Bengali News Headline Generationmogesa/Roberta-amharic-news-sentence-transformerTransformer-based Text Generation for Code-Switched Sepedi-English News

Transformer Based Amharic Headline Generation using Sub-word2Vec Representation

The major component of a news item that aids in giving a quick peek at the article's theme is the he

AfriHG: News headline generation for African Languages

This paper introduces AfriHG -- a news headline generation dataset created by combining from XLSum a

altaseb12/Automatic-Amharic-Text-News-Classification-Using-Fine--Tuned-BERT-Models

Automatic Amharic Text News Classification Using Fine- Tuned BERT Models like XML-R,Afri berta

When Less Is Enough: Context Selection and Prompting Strategies for Bengali News Headline Generation

Large language models (LLMs) have shown strong performance in text generation tasks, yet their effec

mogesa/Roberta-amharic-news-sentence-transformer

Transformer-based Text Generation for Code-Switched Sepedi-English News

Code-switched data is rarely available in written form and this makes the development of la