DeepDub AI Video Translator
An advanced backend system that automatically extracts audio from videos, transcribes the speech, translates it using Large Language Models (LLM), synthesizes new localized audio via TTS, and stitches it back into the original video seamlessly.
Technologies
- Python (FastAPI)
- MoviePy & FFmpeg (Video Processing)
- Google Generative AI (Gemini Flash)
- Google Text-to-Speech (gTTS)
Key Implementations
- Asynchronous background task processing for heavy video workloads
- Simulated terminal UI for real-time frontend status polling
- Automated audio extraction and re-encoding
- Multi-language translation architecture