Logo Lanfrica

Wollo Dialect Identification for Amharic Language Using Machine Learning

Domaine:

natural language processing

Type de record:

paper
Créateur:
Bir
Éditeur:
Zenodo
Hôte:avatar
This thesis presents a study on Wollo dialect identification using machine learning, specifically the Random Forest algorithm. The primary goal of this research is to develop an effective model capable of identifying the Wollo dialect from formal words, contributing to the advancement of natural language processing (NLP) in underrepresented languages. The Random Forest model was chosen due to its robustness in handling complex datasets and its ability to manage both classification and feature selection efficiently. The model was trained and tested on a dataset consisting of text samples, and the final performance achieved an accuracy of 74.38%. The evaluation metrics, including precision, recall, and F1-score, showed balanced performance across the different classes. Despite the promising results, challenges such as class imbalance and limited dataset diversity were identified. This study concludes that while the Random Forest model provides a solid foundation for dialect identification, further improvements in data collection, feature engineering, and model optimization are necessary to enhance the model’s performance for practical applications in dialect classification and language processing tasks. Future work could explore alternative machine learning algorithms and more advanced techniques like deep learning for better accuracy and real-time application potential.