正文 · AI 翻译
我们让 GLM-5.3-Flash 在本地运行速度提升了 3.3 倍!
借助优化的解码以及额外的多 token 预测,本地 GGUF 推理现在快了 1.6-3.4 倍。
通过 Unsloth Desktop 或 llama.cpp,可在 128GB 配置上运行 3-bit 量化。
指南:https://unsloth.ai/docs/models/glm-5.3-flash#faster-inference-and-mtp-support GGUF:https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF
推出 GLM-5.3-Flash - 以极具竞争力的价格提供领先能力 - 原生多模态,具备 1M token 上下文窗口 - 一个以 MIT 许可证发布的 320B-A18B 模型 - 此前以 Ox Alpha 为名预览,完全运行于中国 AI 芯片之上 博客:http://z.ai/blog/glm-5.3-flash...