Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Corpora Preparation and Stopword List Generation for Arabic data in Social Network

Domain:

natural language processing

Record type:

datasetpaper
Creator:
MedYouKor
Publisher:
arXiv
Host:avatar
This paper proposes a methodology to prepare corpora in Arabic language from online social network (OSN) and review site for Sentiment Analysis (SA) task. The paper also proposes a methodology for generating a stopword list from the prepared corpora. The aim of the paper is to investigate the effect of removing stopwords on the SA task. The problem is that the stopwords lists generated before were on Modern Standard Arabic (MSA) which is not the common language used in OSN. We have generated a stopword list of Egyptian dialect and a corpus-based list to be used with the OSN corpora. We compare the efficiency of text classification when using the generated lists along with previously generated lists of MSA and combining the Egyptian dialect list with the MSA list. The text classification was performed using Naïve Bayes and Decision Tree classifiers and two feature selection approaches, unigrams and bigram. The experiments show that the general lists containing the Egyptian dialects words give better performance than using lists of MSA stopwords only. Language Engineering Conference 2014, Cairo, Egypt, 1-3 December 2014

Visit

doi.orgarxiv.org

Tasks

sentiment analysisstopwordstext classification

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similar

Parallel Corpora Preparation for English-Amharic Machine TranslationParallel Corpora Preparation for bi-Directional Amharic-Kistanigna Machine TranslationStopword Lists for African LanguagesStopword Lists for African LanguagesAchieving entrepreneurial intention through entrepreneurial orientation, social network ties, and market intelligence generation perspectivesMohamedHOmar/Somali-Stopword

Parallel Corpora Preparation for English-Amharic Machine Translation

In this paper, we describe the development of an English-Amharic parallel corpus and Machine Translation (MT) experiments conducted on it. Two different tests have been achieved. Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) experiments

Parallel Corpora Preparation for bi-Directional Amharic-Kistanigna Machine Translation

Stopword Lists for African Languages

Some words, like “the” or “and” in English, are used a lot in speech and writing. For most Natural Language Processing applications, you will want to remove these very frequent words. This is usually done using a list of “stopwords” which has been complied by hand.

Stopword Lists for African Languages

Stopword Lists & Frequency Information for 9 African Languages

Achieving entrepreneurial intention through entrepreneurial orientation, social network ties, and market intelligence generation perspectives

Entrepreneurial orientation (ENO), social network ties (SOT) and market intelligence generation (MIT

MohamedHOmar/Somali-Stopword

Somali Stopwords