Introduction
Social media platforms such as Facebook have become central to public discourse,
enabling users to express opinions, share information, and engage in discussions.
For low-resource languages like Somali, analyzing this user-generated content poses
unique challenges due to the lack of linguistic resources and annotated datasets. Sentiment analysis, particularly the distinction between subjective (opinion-based) and objective (fact-based) statements, is crucial for understanding public opinion, detecting
misinformation, and monitoring social trends.
This project aims to build and compare various machine learning and deep learning
models to classify Somali Facebook comments as subjective or objective. The dataset,
collected from Facebook pages discussing Somali politics and society, was manually
labeled and preprocessed. We explore traditional feature-based methods (TF-IDF +
SVM), simple neural architectures (CNN, RNN), and state-of-the-art transformer models (multilingual BERT). Given the inherent class imbalance, we also investigate techniques such as class weighting to improve minority class performance.
The report is organized as follows: Section 2 describes the dataset and its characteristics. Section 3 outlines the preprocessing steps applied to the Somali text.
Section 4 presents exploratory data analysis, including visualizations of label distribution, frequent words, and text length patterns. Section 5 details the methodology for
each model, including hyperparameters and training procedures. Section 6 reports
the results, including accuracy, precision, recall, and F1-score, along with a discussion of model strengths and weaknesses. Finally, Section 7 concludes the study and
suggests directions for future work.
2 Data Collection and Description
The dataset used in this study, final labeled Hsheik data.csv, consists of 3,887
rows of Facebook comments collected from posts related to Somali political figures.
Each record includes the comment text, a la …