Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

KevKibe/Chat-based-Interactions-with-Swahili-Audio-and-Videos

Domaine:

natural language processing

Type de record:

software
Créateur:
Kev
Hôte:
This is an application that enables a chat-based conversation in English with Swahili videos/audio files using Python and deployed as a REST API using Flask containerized using Docker. ## Description This is an application that enables a chat-based conversation in English with Swahili videos/audio files using Python. The application also works for other languages but I am farmiliar with Swahili hence why I picked the language. This is made possible by: - **Pytube** which is a library used to retrieve audio files from Youtube Videos. - **Whisper API** from OpenAI which is used to transcribe Swahili audio files to Swahili text. - **Google Translate API** to translate Swahili text to English. - **Langchain and OpenAI** enables the application to generate dynamic responses and engage in chat-based conversations in English. ## Limitations - Takes a while to transcribe Swahili audio files to text specifically 2mins for a 9min-long audiofile and 13mins for a 46min-long audiofile as shown in the notebook. - Google Translate API for translation is not that accurate. - OpenAI API has a token and request limit and for free trial users. It would be better to use a paid account. - The transcription part of the application requires alot of computation. I used a Tesla P100 GPU and it was running at maximum when transcribing the 46min-long audio ## Usage You can clone this repository and follow these steps or use a notebook as demonstrated here. - Ensure you have Python installed on your machine (version 3.6 or higher). - Install the required dependencies by running the following command: `pip install -r requirements.txt` - Run this command in your terminal `pip install git+github.com ` to get whisper by OpenAI - Install ffmpeg using instructions from this site - Set up environment variables: Create a `.env` file in the root directory of the project and add your OpenAI API key as follows: `OPENAI_API_KEY=your_api_key_here` - Run the app by running the command `python main.py` in your terminal which will prompt you with the youtube URL input and then proceed to ask for a prompt. - The application will transcribe the Swahili audio, …

Visit

github.com

Tasks

automatic speech recognitionspeech processing

Languages

Swahili