One of the stranger assumptions in AI is that the same model architecture should work equally well for thinking and interacting.
A voice model has a weird job, it cannot just produce the right answer. It has to keep up with a person while the sequence keeps getting longer.
In voice AI "speech" may not the right abstraction for the hard part, rather continuous state could be it. Here, Cartesia's founder talking how they came into the space from sequence modeling rather than speech research, which is probably why they focused so heavily on state space models.
--- (Full video on “The Neon Show” YT channel, link in comment)