RiskChainBench:面向混淆平台消息还原与证据溯源网页调查的基准

HuggingFace Daily Papers(社区热门论文)·2026-09-15 08:00·3天前
AI 导读

研究团队推出 RiskChainBench,将 600 个来源会话生成的 3,600 条合成 token 文本还原输入与 600 个对应的人工标注本地网页环境配对,分别评测消息还原与正确路由后的网页调查。

HuggingFace Daily Papers(社区热门论文)
38AI 编辑部评分,满分 100

RiskChainBench:面向混淆平台消息还原与证据溯源网页调查的基准

2026-09-15 08:00· 3天前
AI 导读

研究团队推出 RiskChainBench,将 600 个来源会话生成的 3,600 条合成 token 文本还原输入与 600 个对应的人工标注本地网页环境配对,分别评测消息还原与正确路由后的网页调查。

Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages separately, obscuring how target recovery affects downstream evidence acquisition. We introduce RiskChainBench, pairing 3,600 synthetic token-text restoration inputs from 600 source sessions with 600 corresponding human-labeled local web environments.

A model first restores the message, operational intent, and destination; the same underlying model then acts as a VLM-driven web agent that investigates the correctly associated website and produces a frozen, evidence-cited risk report without message-side semantics or domain-reputation cues. We score restoration and correct-routing web investigation separately and compose them offline by applying the frozen primary-entry prediction as a gate to the same Task 2 result. Human labels determine task correctness, while a fixed multimodal evidence judge assesses faithfulness, sufficiency, completeness, and consistency.

Across ten models, Entry Top-1 ranges from 35.2% to 95.2% and web decision accuracy from 26.3% to 62.8%; the leading systems differ across entry recovery, full reconstruction, website decisions, and fine-grained typing. Execution failures account for 31.9% of web runs, whereas post-decision type errors account for only 0.9%, identifying stable exploration and risk judgment as the principal bottlenecks. We release the benchmark, protocol, and resettable local sandbox.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org