多模态大语言模型(MLLM)通过结合视觉与文本线索在真实场景中进行有依据的预测,然而现有基准测试很少揭示当这些证据来源相互冲突时,模型如何在二者之间进行仲裁。我们引入了 SIGNPOST-Bench,这是一个受控的反事实基准测试,用于评估文本-视觉冲突消解能力。
每个源图像被转换为由原始(Original)、空白(Blank)、相似(Similar)、随机(Random)和对抗(Adversarial)变体组成的反事实五元组。合成、局部化的场景文本干预被设计为保留非文本内容,从而能够对定位性能的变化以及由冲突文本引入的、指向地理目标的定向偏移进行配对测量。
SIGNPOST-Bench 包含来自四个数据集的 5,111 个反事实组和 25,555 个图像变体。我们评估了来自七家提供商的 20 个 MLLM。与原始(Original)图像相比,对抗(Adversarial)变体将中位定位误差从 282 公里提升至 1,347 公里,增加了 4.8 倍。
在可地理编码的对抗样本中,各模型有 6.5%–20.1% 的预测落在注入目标 50 公里以内,并且每个被评估的模型在从空白(Blank)到对抗(Adversarial)的目标距离上均表现出正的平均配对缩减。兼容、无关和冲突的文本替换对模型预测产生不同的影响,而干净输入下的定位性能并不能完全预测对冲突文本的鲁棒性。
这些结果将视觉地理定位确立为场景文本仲裁的连续诊断工具,并为评估 MLLM 如何消解冲突的多模态证据提供了一个受控框架。
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text.
SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text.
These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.