# Unsloth 让 Qwen3.8-Flash 本地推理提速 1.7 倍

- 来源：Unsloth AI (@UnslothAI)
- 发布时间：2026-09-02 22:29
- AIHOT 分数：66
- AIHOT 标记：精选
- AIHOT 链接：https://aihot.news/items/cmtyjea9r0005ro10eeudp015
- 原文链接：https://x.com/UnslothAI/status/2095157074112164072

## 精选理由

原文给出了 MTP 加速的具体倍数、显存门槛和各量化档位内存需求，读者可以据此判断自己的设备能否跑通。

## AI 摘要

Unsloth 通过 MTP（Multi-Token Prediction）让 Qwen3.8-Flash-Next 本地推理提速约 1.3 至 1.7 倍且精度不变，GGUF 在单张 RTX PRO 6000 上可达 170 tokens/s（基线 100 tokens/s）。

## 正文

Qwen3.8-Flash 现在借助 MTP 在本地运行速度提升 1.7 倍！⚡️

GGUF 在 RTX PRO 6000 上可达 170 tokens/s。

MTP 让 Qwen3.8-Flash-Next 的推理速度提升约 1.3-1.7 倍，且精度不变。

GGUF：https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF 指南：https://unsloth.ai/docs/models/qwen3.8-next

### 引用推文

> Unsloth AI：Qwen3.8-Flash 现在可以在本地运行了！🔥 这个 125B 的 MoE 模型表现超过了 Claude-Opus-4.6（Max）。 通过 Unsloth GGUF 在 75GB 内存上运行。 Qwen3.8-Flash-Next 让 CPU 内存 / 统一内存配置也能达到接近 VRAM 的速度。 指南：ht...
