Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Yoshi728-PNG/swahili-customer-service-transcription

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Yos
Hôte:
# English–Swahili Customer Service Transcription and Analysis Dataset A multilingual transcription project demonstrating professional-grade transcript preparation for AI training: speaker diarization, timestamping, preserved English–Swahili code-switching, and Python-based transcript analysis — modeled on real Kenyan customer support interactions. --- ## Project Purpose Many customer support conversations in Kenya mix English and Swahili within the same conversation. AI systems used for speech recognition, chatbot training, call-center analytics, and support automation need accurately transcribed multilingual conversations to train and evaluate against. This project builds a small, high-quality transcription dataset — with documented guidelines and automated quality analysis — that mirrors that real-world need. --- ## Dataset Description Three simulated customer service conversations between an **Agent** and a **Customer**, each covering a distinct real-world scenario: | Interview | Topic | |---|---| | `interview_01` | Online order delivery issue | | `interview_02` | Mobile money (M-Pesa) payment problem | | `interview_03` | Home internet service complaint | Each transcript includes: - Timestamps every 15–20 seconds (`[HH:MM:SS]`) - Consistent speaker labels (`Agent:` / `Customer:`) - Natural English–Swahili code-switching, preserved exactly as spoken - Marked inaudible sections and natural fillers where meaningful --- ## Repository Structure ``` swahili-customer-service-transcription/ ├── audio/ │ └── NOTE.md # Instructions for adding the 3 MP3 recordings ├── transcripts/ │ ├── interview_01.txt │ ├── interview_02.txt │ └── interview_03.txt ├── data/ │ ├── transcript_summary.csv # Generated by transcript_analysis.py │ └── word_count_chart.png # Generated by transcript_analysis.py ├── transcription_guidelines.md ├── transcript_analysis.py ├── requirements.txt ├── README.md └── report.md ``` --- ## Transcription Methodo …

Visit

github.com

Tasks

automatic speech recognitioncode switchingspeech processing

Languages

Swahili

Similaires

Yoshi728-PNG/sentiment-annotation-projectelohimgers-png/nihongo-swahiliabdallah0620885876-png/Learn-Swahilibenny-png/SWAHILI-TTSBenjamin-png/swahili-normalized-corpusBenjamin-png/bert-tokenizer-swahili

Yoshi728-PNG/sentiment-annotation-project

Sentiment annotation quality evaluation on an English-Swahili social media dataset # Sentiment Anno

elohimgers-png/nihongo-swahili

Learn Japanese and Swahili together

abdallah0620885876-png/Learn-Swahili

An online platform for learning Swahili through simple lesson, vocabulary and grammar

benny-png/SWAHILI-TTS

# Swahili Text-to-Speech API A FastAPI-based REST API that provides text-to-speech (TTS) capabilitie

Benjamin-png/swahili-normalized-corpus

Preserved punctuation and special characters Normalized numbers to Swahili words Cleaned formatting

Benjamin-png/bert-tokenizer-swahili