跳到正文
LlamaIndex:产品、工程与评测·· 1 天前AI 评分46

OCR 已死,Agentic OCR 长存:LlamaParse 如何为 AI 智能体解析文档

OCR is Dead. Long Live Agentic OCR.

AI 导读

LlamaIndex 提出 Agentic OCR,即通过 LLM 工具调用循环让解析系统自行决定阅读策略、调用专用模型并基于结果迭代,而非传统 OCR 的单次转换。该系统需满足三点:阅读策略随文档自适应、修正须有证据支撑、执行须有明确时限与目标。LlamaParse 已在 "Agentic" 和 "Agentic Plus" 档位提供上述能力。

正文

OCR Is Dead. Long Live Agentic OCR.

What building LlamaParse has taught us about reading documents for AI agents.

Parsing documents is a hard, long-tail problem. It’s easy to end up with a result that looks plausible (tables, headings, paragraphs, lists), but ends up failing in subtle ways. Parsing failures can dramatically change how downstream readers interpret data (misaligned columns, dropped lines, missing headers/footers).

For hard documents and the edge cases, document parsers that treat the problem as a single step conversion will fail. There needs to be some adaptive, verifiable process that can reason about the text and help tackle those long-tail documents. Building LlamaParse has shown us exactly what that requires in practice, what it looks like, and why its different than traditional OCR.

What is agentic OCR?

Agentic OCR is a version of document parsing in which the system can decide how to read a document, invoke tools or specialized models, and change its approach or iterate based on the results.

Essentially, the word “agentic” means some LLM tool-calling loop. While parsing a document, the system can select a region for closer inspection, request orientation correction, delegate a table to a specialist, or reject an attempted correction.

We can break this type of system down into something that satisfies three requirements: the reading strategy should adapt to the document, corrections should face explicit checks and require evidence, and execution should have clear limits and objectives to satisfy UX.

1. Let the document steer the parsing

Treating all elements on a page the same way makes is what a traditional OCR system would do. While an agentic approach applies extra effort to the content and elements that require it. Plain text with headings is rather simple, and tables and charts require extra processing to parse correctly.

Conventional OCR and other document-processing tools still remain useful within this approach. The value of an agentic system comes from applying its capabilities where they help, with the document itself guiding that choice and reasoning.

In this setting, the quality of the parsed document depends on the model, harness, and on the system's ability to give it an appropriate view of the problem. This is very different from traditional single-shot systems and VLMs that parse documents in a single pass.

2. Self-correction needs evidence

Elements like tables are common failure points when parsing a document, but you still need some way to decide that a parsed table is broken. Furthermore, a correction can improve a table's appearance while making its content less faithful or omitting critical details. Agentic OCR needs to handle both ends of this correction process.

Because of this, validation is a central part of the process. A correction should be backed by evidence that the process improved results and moved to a more correct parsed representation. Without it, another model call simply produces another answer.

3. Systems need clear limits and objectives

A parser's objective is faithful reconstruction. We’ve extensively documented what this looks like through benchmarks like ParseBench and the various ways to measure parse quality. It comes down to faithfully transcribing the entire document, and not excluding “page furniture” like headers/footers, strikethroughs, and more.

A parser may need to interpret layout to keep a footnote attached to the right table or attach a caption to the correct figure instead of transcribing it. Its job is to give the downstream readers a faithful output for reasoning and downstream processing.

Real workloads also impose hard limits. Users need results within an acceptable time and cost before they walk away. A useful system needs a way to stop while still returning a result that is useful and can be acted upon. “Keep trying until it is right” leaves too much unspecified and creates an open-ended task.

What should replace the old OCR contract?

The original promise of OCR was to make documents machine-readable, and AI agents raise the standard for what that means.

Agentic OCR gives document infrastructure responsibility for the entire process: choosing an appropriate reading strategy, obtaining better evidence, checking revisions, and preserving the relationships that make the source meaningful.

That is the direction we are building toward with LlamaParse. As agents take on more consequential work, the systems supplying their evidence need to meet a higher standard than readable output.

Today, LlamaParse provides many of these principles in the “Agentic” and “Agentic Plus” tiers. Our ultimate goal is a parser that is 100% correct, and while that includes a very long tail of documents and error types, we are well on our way.

You can try out LlamaParse today for yourself:

来源:LlamaIndex:产品、工程与评测 · llamaindex.ai