A high-performance, GPU-optimized real-time speech-to-text (STT) streaming server built with WebSocket support for multiple concurrent clients. This project leverages the Kyutai STT model and is optimized for NVIDIA RTX 4090 GPUs, providing low-latency transcription for audio streams.
python websocket speech-recognition speech-to-text stt realtime-audio mochi fastapi realtime-streaming kyutai stt-streaming
-
Updated
Sep 19, 2025 - Python