Georgi Gerganov· @ggerganov · X·· 3 小时前AI 评分59
AI 导读
llama.cpp 可通过 ggml RPC 后端在异构设备间分布式推理,作者 Georgi Gerganov 表示这是高级设置,后续会让普通用户更容易使用。被引用的实测称,MiMo 2.6 Flash 的原生 mxfp4 权重可跨 RTX 6000 GPU 与 M5 笔记本以 40 tokens/sec 通过 10 GbE 运行,llama.cpp 开箱即支持。
正文
llama.cpp can distribute inference on heterogeneous devices through the ggml RPC backend
It's an advanced setting but I think with time we'll make it more accessible to regular users.
It's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the native mxfp4 weights of a state-of-the-art model, on heterogeneous hardware. Supported out of the box in llama.cpp.在 X 查看被引用的帖子
来源:Georgi Gerganov · x.com