Logo Lanfrica

Extracting structured flood information for Nigeria from News Media in GDELT

Domaine:

natural language processingclimate

Type de record:

paper
Créateur:
Tha
Éditeur:
UniUni
Éditeur:
Uni
Hôte:avatar
Floods in African countries are becoming more common and Nigeria, the most populous country in Africa, is no exception. In 2022, Nigeria experienced its worst flooding in a decade, impacting many of its states. These floods, likely exacerbated by climate change altered weather patterns and anthropogenic activities, pose significant risks to Nigerian communities and ecosystems, potentially leading to economic damage and the displacement of populations. The exposure of the population to floods can be approximated using flood-related data from past events. However, this requires access to reliable data and information sources. This thesis presents a NLP-based pipeline for extracting flood-related information from news media in the GDELT database, particularly for data scarce regions such as Nigeria. A custom trained text classification model is utilised to identify newspaper articles relevant to floods from GDELT. Subsequently, a NER model, tailored specifically for Nigeria, is employed to identify place names within these articles. This model is trained thanks to a novel approach developed to fully automate the generation of annotated training data. Ultimately, a rule-based approach is applied to extract quantitative information from news articles. The text classification model attained an F1-score of 80%, while the NER model achieved an accuracy of 86%. Initial spatial analyses for the studied period revealed, that place names primarily from southern Nigerian states, including Anambra, Bayelsa, Delta, Imo and Rivers, were frequently mentioned in flood-related news articles. On a test dataset, the extracted quantitative information achieved a cosine similarity of 78%.

Similaires