SANYAM.SYS☀ DAY SHIFT☾ NIGHT SHIFT
0%boot
All projects

RealTMS

Real-Time Multimodal Translator

Real-time multimodal translator that converts speech to speech across 15+ languages and generates an image from the translated text, chaining four AI models in a single pipeline.

View SourceNovember 2024
RealTMS · SchematicFIG/09
ENहि15+ languages · 4-stage pipeline

Speech translation usually stops at text. RealTMS goes speech to speech across 15+ languages and adds a generated image of what was said, all from one Gradio interface.

Gradio Web Interface
OpenAI Whisper (ASR)
SeamlessM4T (Neural Translation)
gTTS (Speech Synthesis)
Stable Diffusion (Image Generation)
  • Speech-to-speech translation across 15+ language pairs
  • Whisper handles speech recognition with good noise tolerance
  • SeamlessM4T performs the multilingual neural machine translation
  • Stable Diffusion generates a contextual image from the translated text
  • Upload or record audio directly in the Gradio interface
  • CPU-only fallback mode: a GPU speeds things up but is not required
  • Four-stage pipeline: Whisper for speech recognition, SeamlessM4T for translation, gTTS for speech synthesis, and Stable Diffusion for image generation.
  • Generated images hit 85% visual relevance against the translated text.