# Unsloth 发布 GLM-5.3-Flash GGUF 量化版本地运行方案

- 来源：Unsloth AI (@UnslothAI)
- 发布时间：2026-08-27 22:43
- AIHOT 分数：71
- AIHOT 标记：精选
- AIHOT 链接：https://aihot.news/items/cmtyjea9r0007ro10vl9cld30
- 原文链接：https://x.com/UnslothAI/status/2092986464196002094

## 精选理由

原文给出了各量化档位的内存需求和精度保留数据，读者可据此判断在 128GB 设备上本地运行该模型的可行性。

## AI 摘要

Unsloth 发布 GLM-5.3-Flash（ox-alpha）的 GGUF 量化版本，可在 128GB RAM 设备上以 3-bit 运行。该模型为 Z.ai 的 320B-A18B 开源多模态模型，MIT License 发布，原文称其在 DeepSWE、编码和智能体基准上媲美 Claude Opus 4.8。

## 正文

GLM-5.3-Flash 现在可以在本地运行了！✨

通过 Unsloth GGUF 在 128GB 内存上以 3-bit 量化运行。

GLM-5.3-Flash（ox-alpha）在 DeepSWE、编程和智能体基准测试上与 Claude Opus 4.8 不相上下。

指南：https://unsloth.ai/docs/models/glm-5.3-flash GGUF：https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF

### 引用推文

> Z.ai：隆重推出 GLM-5.3-Flash - 以极具竞争力的价格提供领先能力 - 原生多模态，具备 1M-token 上下文窗口 - 一个以 MIT 许可证发布的 320B-A18B 模型 - 此前以 Ox Alpha 为名预览，完全运行于中国 AI 芯片之上 博客：http://z.ai/blog/glm-5.3-fla...
