Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not whether evidence remains traceable through rationale, confidence, and final action. We introduce E2A-Bench, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors. E2A-Bench evaluates grounding, reasoning-action consistency, evidence-confidence calibration, and directional coverage through UCR, RCI, ECI, and NDR, where NDR measures coverage-aware evidence-to-action reliability rather than realized trading performance. Evaluating 20 VLMs reveals three failures hidden by scalar hallucination scores: the lowest-UCR model ranks near the bottom by NDR due to only 6.4% directional coverage; oracle-aided verification reduces unsupported claims but can collapse coverage; and financial fine-tuning amplifies the BUY:SELL ratio by factors of 4.21 to 4.68 across strict base-fine-tuned pairs. These results show that financial VLM evaluation should trace the full evidence-to-action chain rather than rely on a single hallucination score. Code and data: https://github.com/wanng-ide/E2A-Bench
E2A-Bench:金融图表推理中的证据到行动可靠性基准
AI 导读
研究者推出 E2A-Bench,一个包含 969 条查询的金融图表推理基准,基于 323 只 HS300 成分股和三种输入模态构建,用 UCR、RCI、ECI、NDR 四项指标评估证据溯源、推理-行动一致性、证据-置信度校准与方向覆盖。
HuggingFace Daily Papers(社区热门论文)
40
AI 编辑部评分,满分 100E2A-Bench:金融图表推理中的证据到行动可靠性基准
研究者推出 E2A-Bench,一个包含 969 条查询的金融图表推理基准,基于 323 只 HS300 成分股和三种输入模态构建,用 UCR、RCI、ECI、NDR 四项指标评估证据溯源、推理-行动一致性、证据-置信度校准与方向覆盖。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org