Cartesia 创始人谈语音 AI 为何押注状态空间模型

Rohan Paul · @rohanpaul_ai · X·2026-09-22 00:08·1小时前
AI 导读

Cartesia 创始人认为语音模型的难点不在"语音"本身,而在于持续状态:它不能只给出正确答案,还要在序列不断变长时跟上对话节奏。他提到团队是从序列建模而非语音研究进入这一领域,这也是其重点投入状态空间模型的原因。相关完整视频发布于 YouTube 频道 The Neon Show。

Rohan Paul@rohanpaul_ai
33AI 编辑部评分,满分 100

Cartesia 创始人谈语音 AI 为何押注状态空间模型

2026-09-22 00:08· 1小时前
AI 导读

Cartesia 创始人认为语音模型的难点不在"语音"本身,而在于持续状态:它不能只给出正确答案,还要在序列不断变长时跟上对话节奏。他提到团队是从序列建模而非语音研究进入这一领域,这也是其重点投入状态空间模型的原因。相关完整视频发布于 YouTube 频道 The Neon Show。

One of the stranger assumptions in AI is that the same model architecture should work equally well for thinking and interacting.

A voice model has a weird job, it cannot just produce the right answer. It has to keep up with a person while the sequence keeps getting longer.

In voice AI "speech" may not the right abstraction for the hard part, rather continuous state could be it. Here, Cartesia's founder talking how they came into the space from sequence modeling rather than speech research, which is probably why they focused so heavily on state space models.

--- (Full video on “The Neon Show” YT channel, link in comment)

来源:Rohan Paul· x.com