🎙️ Real-time interpretation, powered locally by VoxCPM2.
Developer @HenryZ30734018 built VoxWeft, an open-source simultaneous interpretation system for Apple Silicon. It uses an MLX implementation of VoxCPM2 to turn live speech into translated speech on-device, keeping audio private and responsive.
✨ Highlights: ⚡ VoxCPM2 streams first audio in ~170 ms on an M5 MacBook 🌍 Generates speech across 30 languages, supporting direct language-pair interpretation without a pivot 🗣️ Clones a target voice from ~5 seconds of reference audio for a consistent interpreted voice 💻 Runs in 4-bit quantization on MLX, making low-latency local speech generation practical on Apple Silicon
VoxWeft shows how VoxCPM2 can become the speech layer of a full real-time application: not just producing audio, but enabling private, multilingual interaction that stays on-device.
Try VoxCPM2 and see what you can build with it!
🔗GitHub:http://github.com/HenryZ838978/VoxWeft 🔗Devlog:http://github.com/HenryZ838978/VoxWeft/blob/main/docs/DEVLOG.md 🤗 VoxCPM2:http://huggingface.co/openbmb/VoxCPM2