🤖 What if an AI agent could keep seeing, listening, talking, and working — all at the same time?
Developer @speechjsp built Gander, a multimodal duplex interaction agent that combines continuous audio-visual perception, real-time conversation, and asynchronous agent execution.
Gander’s Cerebellum is fine-tuned from MiniCPM-o 4.5, with additional training and system-level improvements for native duplex interaction.
Instead of treating voice interaction and agent execution as separate steps, Gander brings them together in one continuous interaction loop.