跳到正文
原文
AK· @_akhaliq · X·· 2 小时前AI 评分48
AI 导读

Audio8 ASR Infinite 是一款原生流式语音识别模型,每秒解码 12.5 次,通过滚动 KV Cache 在 24/7 持续运行下保持内存与延迟恒定。模型每个时钟步输出一个文本 token(12.5/8.3/6.25 次决策/秒),平衡感知粒度与资源开销。HuggingChat 的 ML intern 已搭建 Gradio 工作流供试用。

正文

Audio8 ASR Infinite

streaming speech recognition model

the native streaming architecture decodes 12.5 times per second

a rolling KV Cache keeps both memory and latency constant, even in 24/7 operation

one text token per clock step (12.5 / 8.3 / 6.25 decisions per second), balancing perception granularity and resource cost

ML intern in huggingchat setup a gradio workflow to try it out: https://huggingface.co/spaces/akhaliq/audio8-asr-workflow

huggingchat: https://huggingface.co/chat/

model: https://huggingface.co/Edge0/Audio8-ASR-Infinite

来源:AK · x.com