Logo Lanfrica

Social Media Prediction Challenge

Domain:

natural language processing

Record type:

dataset

Predict which tweets from major African companies will receive the most retweets
The data has been split into a test and training set.
train.json (zipped) is the dataset that you will use to train your model. This dataset includes about 2,400 consecutive tweets from each of the companies listed below, for a total of 96,562 tweets.
test_questions.json (zipped) is the dataset to which you will apply your model to test how well it performs. Use your model and this dataset to predict the number of retweets a tweet will receive. The test set are the consecutive tweets that followed the first tweets provided in the training sets. There are a maximum of 800 tweets per company in this test set. This dataset includes the same fields as train.json except for the retweet_count and favorite_count variables.
sample_submission.csv is a table to provide an example of what your submission file should look like.
Notes on the data: This data was downloaded from Twitter on 23 August 2018. So represents the retweets and favorites at that point in time.
Variables in train.json and test_questions.json are as described in the twitter documentation:
Tweet Object - developer.twitter.com
User Object - developer.twitter.com
Entities Object- developer.twitter.com
GeoObject - developer.twitter.com
Companies included in this dataset:
Nigeria
Zenith Bank
First Bank Nigeria
Guaranty Trust Bank
Access Bank
Diamond Bank
Ecobank
MTN
Airtel
GloMobile
Ghana
Barclays
Fidelity
Ecobank
Access
Ghana commercial bank
MTN
Vodafone
Airtel Tigo
South Africa
Standard Bank
ABSA-Barclays
FNB
Nedbank
Capitec
Vodacom
MTN
Cell C
Telkom
Kenya
Equity
Kenya Commercial Bank
Co-operative Bank
Standard Chartered
Safaricom
Airtel
Telkom
Uganda
Stanbic
DFCU
Standard Chartered
MTN
Airtel