WeVisDoc:从覆盖到能力的两阶段文档解析框架,WeVisDoc-4B 在 OmniDocBench v1.6 上达 95.38

HuggingFace Daily Papers(社区热门论文)·2026-09-17 08:00·1天前
AI 导读

WeVisDoc 是一个两阶段数据中心化端到端文档解析框架:Stage I 通过异构数据与结构保持退化合成扩展语义、结构和外观覆盖,Stage II 用留出探针在固定视觉-结构聚类中定位残余错误,指导定向数据构建与目标 token 预算再分配。

HuggingFace Daily Papers(社区热门论文)
36AI 编辑部评分,满分 100

WeVisDoc:从覆盖到能力的两阶段文档解析框架,WeVisDoc-4B 在 OmniDocBench v1.6 上达 95.38

2026-09-17 08:00· 1天前
AI 导读

WeVisDoc 是一个两阶段数据中心化端到端文档解析框架:Stage I 通过异构数据与结构保持退化合成扩展语义、结构和外观覆盖,Stage II 用留出探针在固定视觉-结构聚类中定位残余错误,指导定向数据构建与目标 token 预算再分配。

Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser's remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis.

Stage II uses a held-out probe to measure the Stage I parser's residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org