跳到正文
Hacker News:AI 热帖· fibo·· 5 小时前AI 评分64

Redis 作者推出 ds4 本地推理引擎,可跑 DeepSeek V4、GLM 5.x 和 Qwen3.8

From the creator of Redis; run LLM locally with ds4

AI 导读

ds4 是一个面向高内存 Mac、CUDA 和 ROCm 机器的窄口径 C 语言推理引擎,采用 MIT 许可证开源,支持 DeepSeek V4/V4.1 Flash、GLM 5.x 和 Qwen3.8 Flash Next 的文本与视觉模型。

正文

DS4 · LOCAL FRONTIER INFERENCE

Run frontier open weights locally with ds4.

DwarfStar 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.

SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE · C / METAL / CUDA / ROCM · QWEN ON 64GB

ds4 · local session

PRINCIPLE · LOCAL MODEL STACK

PHASE 1 · THE GIANT

A 284-billion-parameter star

DeepSeek V4 Flash is a large mixture-of-experts model. The usual path is remote serving; ds4 starts from the opposite constraint.

PHASE 2 · THE COLLAPSE

Compressed, not lobotomized

Asymmetric quantization targets the routed experts while preserving critical paths. The model becomes practical on high-memory machines.

PHASE 3 · THE DWARF STAR

Dense, resident, yours

The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.

How the collapse works →

SCROLL ▾

DESIGN CHOICES

Local frontier inference, narrow on purpose

Not a generic GGUF runner. ds4 follows a small, opportunistic set of model families and validates each supported layout end to end.

CORE 01

Asymmetric 2-bit quantization

Compress the routed experts, keep critical shared paths precise. That is how the supported routed-MoE builds fit their target machines.

CORE 02

KV cache as a disk citizen

Save long prefixes to SSD and resume by prompt hash. Restarts do not have to mean full re-prefill.

CORE 03

One engine, three interfaces

Use ./ds4 for chat, ./ds4-server for local APIs and ./ds4-agent for persistent coding sessions.

  • SSD STREAMING
  • TENSOR PARALLELISM
  • SESSION BATCHING
  • DSPARK + MTP
  • VISION INPUT

ARCHITECTURE

How the ds4 stack fits together

Project GGUFs, a self-contained engine and agent-facing interfaces, checked against official model outputs.

RUNTIME MAP · SIMPLIFIED. SEE ARCHITECTURE NOTES FOR THE FULL DRAWING.

RUN IT

Run ds4 in three steps

Download the project GGUF, build for your backend, then start the CLI or server. Generic GGUF files are not the target.

STEP 1 · FETCH THE WEIGHTS

ds4 · zsh

$ git clone https://github.com/antirez/ds4
$ cd ds4 && ./download_model.sh ds4f-q2

STEP 2 · BUILD FOR YOUR BACKEND

ds4 · zsh

$ make
$ make cuda-spark

STEP 3 · TALK TO IT

ds4 · zsh

$ ./ds4
$ ./ds4-server --ctx 100000

FIT CHECK

ds4 hardware fit: local, streamed and distributed

Pick your platform and memory: get a conservative starting path and understand which execution modes apply.

✓ Runs well

V4 Flash Q2 is the baseline. At 128 GB, GLM 5.3 Q2 and Qwen Q4 also fit; V4.1 Q2 streams from SSD.

./download_model.sh ds4f-q2 && make

REF · M5 MAX 128GB · 32K CTX: 34.4 T/S GEN · 557 T/S PREFILL

Estimates from the ds4 benchmark table. Full guide in Hardware and Installation.

BENCHMARKS

ds4 benchmarks: prefill and generation

Reference rows from upstream. Read prefill and generation separately, especially for long-context agent workloads.

Machine Context Prefill t/s Generation t/s
M5 Max, 128 GB q2 · 2,048 tok 790.2 39.4
M5 Max, 128 GB q2 · 65,536 tok 398.5 27.6
DGX Spark, 128 GB q2 · 2,048 tok 825.8 18.1
DGX Spark, 128 GB q2 · 65,536 tok 823.0 13.8

All benchmarks →

API & AGENTS

Use ds4 from Codex, Claude Code and OpenCode

ds4-server speaks OpenAI and Anthropic-style APIs, so local coding agents can connect to your own machine with a base URL.

Own your local AI inference.

Start with the quickstart, check the hardware matrix, then connect your editor, agent or API client to the local server.

来源:Hacker News:AI 热帖 · dwarfstar.sh