inclusionAI 发布 SingProbe:gemma-4-26B-A4B-it 流式安全探针

蚂蚁 inclusionAI:HuggingFace 新模型·2026-09-14 10:55·18小时前
AI 导读

蚂蚁 inclusionAI 发布 SingProbe,一个构建在 google/gemma-4-26B-A4B-it 上的流式安全探针,复用基座模型隐藏状态逐 token 打分,探针参数仅 5.67M,解码开销低于 0.5%。其流式安全 R-AUC/T-AUC 达 0.9740/0.9175,幻觉检测 AUC 0.7952,已支持 SGLang 与 vLLM 集成。

蚂蚁 inclusionAI:HuggingFace 新模型
44AI 编辑部评分,满分 100

inclusionAI 发布 SingProbe:gemma-4-26B-A4B-it 流式安全探针

2026-09-14 10:55· 18小时前
AI 导读

蚂蚁 inclusionAI 发布 SingProbe,一个构建在 google/gemma-4-26B-A4B-it 上的流式安全探针,复用基座模型隐藏状态逐 token 打分,探针参数仅 5.67M,解码开销低于 0.5%。其流式安全 R-AUC/T-AUC 达 0.9740/0.9175,幻觉检测 AUC 0.7952,已支持 SGLang 与 vLLM 集成。

Model Description

SingProbe is an intrinsic streaming guardrail built on google/gemma-4-26B-A4B-it. Rather than running a separate safety model, this lightweight probe reuses the base model's hidden states during generation to score, at every token, query intent, response unsafety, and hallucination risk. It adds less than 0.5% decode-time overhead.

Base model Probe parameters Tapped layers Outputs
inclusionAI/gemma-4-26B-A4B-it-singprobe 5.67M [8, 18, 28] 8 intents + unsafe + hallucination

See the technical report for methodology and complete results. Training codes are available at inclusionAI/SingProbe.

Evaluation

Higher is better for every metric. Results are averages over the benchmark suites specified below.

Task Metric gemma-4-26B-A4B-it-singprobe Reference baseline
Query intent classification (6 benchmarks) F1 0.8439 YuFeng-XGuard-Reason-8B: 0.8714
Response safety classification (8 benchmarks) F1 0.8647 Qwen3Guard-Gen-8B-strict: 0.8604
Streaming safety (3 benchmarks) R-AUC / T-AUC 0.9740 / 0.9175 Qwen3Guard-Stream-8B-strict: 0.9640 / 0.8893
Hallucination detection (6 benchmarks) AUC 0.7952 DRIFT: 0.8000
Deployment characteristic Result
Benign-response false-positive rate 0.04% average across 5 datasets
Decode overhead < 0.5%

Quick Start

SingProbe is supported through the SGLang integration branch or vLLM integration branch. Load the probe by its Hugging Face ID at server launch:

python -m sglang.launch_server \
  --model-path google/gemma-4-26B-A4B-it \
  --probe-ckpt inclusionAI/gemma-4-26B-A4B-it-singprobe \
  --port 30000

The integrations return one score dictionary per generated token (label_0label_9). Use the exact base-model/probe pair: google/gemma-4-26B-A4B-it with this checkpoint.

Citation

@article{singteam2026singprobe,
  title = {SingProbe Technical Report},
  author = {Sing Team},
  journal = {arXiv preprint arXiv:2608.30703},
  year = {2026},
}

来源:蚂蚁 inclusionAI:HuggingFace 新模型· huggingface.co