Logo Lanfrica

C-Ronny/Vision-To_Speech

Domain:

natural language processing
Creator:
C-R
Host:
VTTS Project that uses a pre-trained models, takes a video input and produces an output of a speech in a specified African language. # Vision-to-Speech Project This project utilizes a pre-trained model to derive actions from a video input and output speech descriptions of these actions. The pipeline involves video processing, action description generation, semantic similarity filtering to ensure unique descriptions, and speech synthesis. ### Overview The Vision-to-Speech project aims to convert video inputs into spoken descriptions. It processes each frame of the video, generates descriptive captions using a pre-trained BLIP model, filters out redundant descriptions, translates the text if necessary, and converts the final description to speech. ### Features Frame to Text: Converts video frames to descriptive text using a pre-trained BLIP model. Unique Meaning Extraction: Filters out similar text descriptions to retain only unique meanings using semantic similarity. Translation: Translates text descriptions into a specified language using Google Translate. Text to Speech: Converts the text descriptions into speech using Google Text-to-Speech (gTTS). ### Requirements Python 3.7 or higher Libraries: - os - cv2 (OpenCV) - PIL (Pillow) - transformers - sentence_transformers - googletrans - gtts - espeak - speak ### Setup and Installation ##### Clone the repository: ##### Install the required libraries: - pip install opencv-python - pip install Pillow - pip install transformers - pip install sentence-transformers - // pip install googletrans==4.0.0-rc1 - pip install googletrans==4.0.2 - pip install gtts ### Usage Prepare your video file and place it in the project directory. Update the video_path variable in the script with the path to your video file. Run the script: app.py ### Testing the Model To test the accuracy of the model, you can manually compare the generated descriptions with the actual video content. ### Hosting the Application (Using Streamlit) Run the program on your device terminal and open the location of your app.py file and run the code: streamlit run app.py ### Link to …