作者在 Cartesia playground 用约 15 秒音频克隆自己的声音,几秒内完成,并让克隆声音说出他不会讲的日语。该功能基于最新语音模型 Sonic 3.6,支持 44 种语言,作者认为可用于让内容覆盖更多语言受众。
This is impressive!
I cloned my voice in the @cartesia playground from a ~15-second clip. The clone was ready in a couple of seconds.
Then I had it speak Japanese, a language I don't speak.
My clone now says in Japanese that I can find good papers by combining AI with my experience reviewing papers.
I am impressed by how good this sounds.
This runs on Sonic 3.6, their latest voice model, which supports 44 languages. You can do the same with any of them.
This tech is getting really good, and I think it's worth exploring to make your content more accessible to more people in different languages.
Try it here: http://play.cartesia.ai
More info about their model here: https://www.cartesia.ai/blog/sonic-3.6
来源:elvis · x.com