Logo Lanfrica

Lucamiras/kaggle_rwanda_emissions

Domaine:

climateenvironment and energy

Type de record:

dataset
Créateur:
Luc
Hôte:
Kaggle competition to predict CO2 emissions in Rwanda # kaggle_rwanda_emissions Kaggle competition to predict CO2 emissions in Rwanda ## Goal - Create a model to predict future CO2 emissions using open source data from Rwanda # EDA Large outliers where massively more emissions were recorded could skew the model Visualizing where most carbon emissions occured by mapping mean emissions to longitude / latitude coordinates. On the first map, color and size both correlate to the amount of emissions over the entire time period. On the second map, I used plotly express to better visualize the emissions across the country. Visualizing emissions over time reveals a somewhat cyclical pattern where emissions are highest in May and November. ## Feature engineering / data cleaning - The dataset contained some missing values. Where values were missing for most observations (UvAerosol*), features were dropped. For others where only a fraction of observations were missing, I used SkLearn's SimpleImputer to impute missing values with the mean. - Dropped non-numerical features such as the ID_LAT_LON_YEAR_WEEK feature, which is just an ID generated from other features, as well as the year_week feature which was replaced by a datetime feature ## Model building I tried three different models (Decision Tree, Linear Regression, Random Forest Regressor). I didn't expect the linear model to do well due to the nature of the data, which was confirmed when running the model. I then repeated this three times: Once for all the data, once for data without outliers beyond 99.5% of range of values, and one only containing only the two important features (longitude and latitude) as well as year and week. | |importance %| |-|-| |longitude|0.46| |latitude|0.44| | ... | ... | ## Model performance The Random Forest outperformed the single Decision Tree by a small margin in every version. I used root mean squared error for evaluation because the rmsq is easily understandable and comparable to the target variable. | | Without outliers | With ou …