Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Deep learning approach for Amharic sentiment analysis using scraped social media data

Domaine:

natural language processing

Type de record:

paper
Créateur:
Yes
Éditeur:
Spr
Hôte:
Abstract Deep learning has emerged as a powerful machine learning technique that learns multiple layers of representations or features of the data and produces state-of-the-art prediction results. Along with the success of deep learning in many other application domains, deep learning is also popularly used in sentiment analysis in recent years. Social media is now playing a vital role in influencing people's sentiment in favors or against a government or an organization. Therefore, to understand the sentiment of any posting in social media an efficient method is an ultimate necessity. We have analyzed some Facebook postings to understand socio-political sentiments. Among this broad scope we have done on this research applying the state of the art in sentiment analysis on Amharic language using deep learning approach in socio-political domain. The further preparation of the dataset on the other domain enhance the language so on this research try to show the data extracted from Fana broadcasting corporation official Facebook page using Graph Application interface of Facebook social media on immigration, war and public relation issues and prepare the data for further preprocessing. After collecting thus data from the FBC using post_id all preprocessing steps tokenization, stop word removal, stemming of the sentence are undertaken. The manual annotation of the sentence extracted data contain both the text file and Emoji are annotated using linguistic experts in seven class, positive, very positive, extremely positive, neutral, negative, very negative and extremely negative class by considering the effect of most common Emoji. Using the Scikit-learn feature extraction classes, Count Vectorizer and TF-IDF Vectorizer the researcher build the feature extraction method to use our custom vocabulary. We test both feature extractors and find that our model performs better with Count Victories as it offers a simple representation of our data. To evaluate the performances of the systems we have collected 11,000 reviews from immigration, war and public relation domains. Then we evaluate our system training and validation accuracy using three experiments by changing training and testing split 90%,10%:80%,20%:70%,30%, the size of the dataset, the number of epochs and network layers. Accordingly, on the first experiment register 90.1% average training accuracy and 90.1% average validation accuracy performances were obtained by the first method. The second method achieves an average training of 82.4% and an average validation accuracy of 40% performances obtained. The third experiment conducted by increasing the number of data set 1600 and five network layers we get 70.1 training accuracy and 40.1 validation accuracy. The results show that this study is promising. A preprint of this research article is previously been published in researchgate by the author [1].

Visit

doi.org

Tasks

sentiment analysistext classification

Languages

Amharic

Licenses

https://creativecommons.org/licenses/by/4.0/