Logo Lanfrica

ermijeremy/Week-2-Challenge-Document-Customer-Experience-Analytics-for-Fintech-Apps

Domaine:

natural language processing

Type de record:

project
Créateur:
erm
Hôte:
This project aims on analyzing the data scraped from google play for three Ethiopian banks ( BOA, Commercial Bank of Ethiopia, Dashen Bank ) to improve their mobile apps # 📱 Fintech App Review Analysis – Week 2 This project is part of the **10 Academy AI Mastery Week 2 Challenge**, where the goal is to analyze Google Play Store reviews for three major Ethiopian banking apps — **CBE**, **BOA**, and **Dashen Bank** — to extract insights that can improve customer experience. --- ## 🎯 Project Objective - Scrape user reviews from the Play Store - Clean and preprocess the review text - Perform sentiment analysis and thematic grouping - Store the processed data in a structured Oracle database - Visualize insights and provide recommendations to each bank --- ## 🛠️ Tools & Libraries - `google-play-scraper` – for scraping app reviews - `pandas`, `numpy` – data processing - `matplotlib`, `seaborn`, `wordcloud` – data visualization - `TextBlob`, `NLTK`, `VADER`, `TextBlob`, `TF-IDF`, `SpaCy` – sentiment & theme extraction - `cx_Oracle` – connecting to Oracle XE database --- ## 🧪 Methodology ### 🔹 Data Scraping We used the `google-play-scraper` Python library to extract user reviews for three Ethiopian banking apps: **CBE**, **BOA**, and **Dashen Bank**. For each app, over 400 reviews were collected to ensure diversity and volume. Key fields extracted: - `review`: User-written review content - `rating`: Integer rating (1 to 5) - `date`: Date the review was posted - `source`: Set to "Google Play" - `bank`: One of "CBE", "BOA", or "Dashen" The scraping was done in batches and saved as raw `.csv` files per app in the `data/raw_reviews/` folder. --- ### 🔹 Data Cleaning & Preprocessing After scraping, each dataset underwent preprocessing to ensure quality and consistency. Steps included: - **Removing Duplicates**: Eliminated repeated reviews based on text content. - **Standardizing Text**: Converted reviews to lowercase, removed URLs, non-ASCII characters, and excess whitespace. - **Filtering Noise**: Removed reviews with fewer than 3 words or missing key fields (e.g., no rating or date). - **Parsing Dates**: Standardized review dates …