E2A-Bench:金融图表推理中的证据到行动可靠性基准

HuggingFace Daily Papers(社区热门论文)·2026-09-13 08:00·3天前
AI 导读

研究者推出 E2A-Bench,一个包含 969 条查询的金融图表推理基准,基于 323 只 HS300 成分股和三种输入模态构建,用 UCR、RCI、ECI、NDR 四项指标评估证据溯源、推理-行动一致性、证据-置信度校准与方向覆盖。

HuggingFace Daily Papers(社区热门论文)
40AI 编辑部评分,满分 100

E2A-Bench:金融图表推理中的证据到行动可靠性基准

2026-09-13 08:00· 3天前
AI 导读

研究者推出 E2A-Bench,一个包含 969 条查询的金融图表推理基准,基于 323 只 HS300 成分股和三种输入模态构建,用 UCR、RCI、ECI、NDR 四项指标评估证据溯源、推理-行动一致性、证据-置信度校准与方向覆盖。

Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not whether evidence remains traceable through rationale, confidence, and final action. We introduce E2A-Bench, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors. E2A-Bench evaluates grounding, reasoning-action consistency, evidence-confidence calibration, and directional coverage through UCR, RCI, ECI, and NDR, where NDR measures coverage-aware evidence-to-action reliability rather than realized trading performance. Evaluating 20 VLMs reveals three failures hidden by scalar hallucination scores: the lowest-UCR model ranks near the bottom by NDR due to only 6.4% directional coverage; oracle-aided verification reduces unsupported claims but can collapse coverage; and financial fine-tuning amplifies the BUY:SELL ratio by factors of 4.21 to 4.68 across strict base-fine-tuned pairs. These results show that financial VLM evaluation should trace the full evidence-to-action chain rather than rely on a single hallucination score. Code and data: https://github.com/wanng-ide/E2A-Bench

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org