社区为 MiniCPM5-2B 构建了 EXL3 4-bit 量化版本,量化后模型权重仅 1.61 GB,在 NVIDIA Tesla T4 上实测约 68–70 tokens/s。该版本采用 4.0 bpw EXL3 量化,支持通过 ExLlamaV3 和 TabbyAPI 进行本地推理。
⚡ Making MiniCPM5-2B even more lightweight for local inference!
A community-built EXL3 4-bit quantization of MiniCPM5-2B brings the quantized model weights down to just 1.61 GB.
✨ Highlights:
⚡ 4.0 bpw EXL3 quantization for a smaller model footprint
🚀 ~68–70 tokens/s verified on an NVIDIA Tesla T4
💾 Just 1.61 GB for the quantized model weights
🛠️ Supports local inference with ExLlamaV3 and TabbyAPI
A great community contribution showing how MiniCPM5-2B can be optimized for more lightweight and accessible local AI deployments.
🤗 Model: http://huggingface.co/ewin-reg/MiniCPM5-2B-EXL3-Quantized
🤗 Base model: http://huggingface.co/openbmb/MiniCPM5-2B
来源:OpenBMB · x.com