LlamaParse 推出 Enriched Forms 表单解析,指出 VLM 难以正确读取表单的原因
Why VLMs Can't Read a Form
LlamaIndex 发布 LlamaParse 的 Enriched Forms 功能,将表单解析为字段与分组的 JSON 树,附带 bounding box 定位和归属信息,供下游审计验证。
A form is one of the most critical types of documents for a business since it carries valuable input data from clients and customers. On top of the complexities that come with processing a large amount of text, including cost, performance, and latency, a form presents a unique set of challenges for any document processing pipeline. Forms carry more structure than standard Markdown can represent, including textboxes, checkboxes, and labels that own the control next to them. So the system has to be designed around those properties specifically. The output has to stay consistent across documents of the same form type, while still absorbing the variations that show up in the same form from one year to the next. Hallucination hurts more here than almost anywhere else. A form has a precise structure, and getting it slightly wrong changes the meaning completely. Whether a checkbox is checked or unchecked can flip the entire downstream action an agent takes. In this post, we walk through why forms need special processing, how LlamaParse represents a form as JSON, and the failure modes we see when asking a VLM to do the same job.
Why forms need a purpose-built parser
There are many ways to parse a form, but they vary a lot in price and performance. While a VLM is a reasonable tool to experiment with (just pass in a page screenshot and ask for an output), it comes with challenges in terms of cost, reliability, and performance. A VLM is optimized for general reasoning, not for document or form processing, and our own results have shown that increased reasoning doesn’t really translate to better parse results.
LlamaParse is built specifically for document and form understanding, and it beats a general VLM at a fraction of the cost. Under the hood, optimized pipelines for form content extraction and bounding box detection are tuned and combined for optimal results.
Representing a form’s structure
A form can be represented as a tree of fields and sections. Fields may have an id , a label , and a value . Their field property identifies the input type, while a section groups related fields in items . Below are examples of elements in our forms schema:
- id: The short designator printed on the form, such as
1,12aore. - label: The text describing what that field is.
- field: The type of a field
- checkbox: A box with a boolean value:
truewhen checked andfalsewhen unchecked. - text: A field for entered text, including numbers, dates and addresses.
- signature: A signature box, either signed or unsigned
- checkbox: A box with a boolean value:
- value: Value of the field. Text for a text field; true/false for a checkbox or signature
The form object can be represented via the simple Pydantic classes below. The structure can have an arbitrary level of nesting, according to the real complexity of the form.
View a minimal schema as Pydantic classes:
from typing import Literal, Optional, Union
from pydantic import BaseModel, Field
class FormField(BaseModel):
"""One entry on the form: a text box, a checkbox, a signature, or a group of choices."""
field: Literal["text", "checkbox", ...]
id: Optional[str] = Field(None, description="Designator printed on the form, e.g. '1', '12a', 'e'")
label: Optional[str] = Field(None, description="Caption printed next to the field")
value: Optional[Union[str, bool]] = Field(
None, description="Verbatim text for a text field; true/false for a checkbox or signature")
class FormSection(BaseModel):
"""A printed grouping of fields, such as 'Part III' or 'Sign Here'."""
type: Literal["section"] = "section"
id: Optional[str] = None
label: Optional[str] = None
items: list["FormNode"]
FormNode = Union[FormField, FormSection] Figure 1 shows how those pieces fit together on a W-2. Box 1 becomes a text field with the id 1 , the label Wages, tips, other compensation and the value 230303.03 . Box 13 stays one group with three checkbox options. Box 9 is empty, but it is still present in the output.
View JSON output
[
{
"type": "field",
"field": "text",
"id": "a",
"label": "Employee's social security number",
"value": "827-37-3673"
},
{
"type": "field",
"field": "text",
"id": "1",
"label": "Wages, tips, other compensation",
"value": "230303.03"
},
{
"type": "field",
"field": "text",
"id": "9",
"isEmpty": true
},
{
"type": "field",
"field": "multi_select",
"id": "13",
"valueItems": [
{
"type": "field",
"field": "checkbox",
"label": "Statutory employee",
"value": false
},
{
"type": "field",
"field": "checkbox",
"label": "Retirement plan",
"value": false
},
{
"type": "field",
"field": "checkbox",
"label": "Third-party sick pay",
"value": true
}
]
},
{
"type": "section",
"id": "12a",
"items": [
{
"type": "field",
"field": "text",
"label": "Code",
"value": "H"
},
{
"type": "field",
"field": "text",
"value": "8699"
}
]
}
] Figure 1. W-2 source and form representation. The source page appears on the left; a form representation json array appears on the right.
While getting the individual fields right is the first step, the output also has to preserve how they relate to one another. Figure 2 shows two printed sections with separate “Date” fields. A flat list of elements loses which section each field came from, so the two Date fields would become indistinguishable. LlamaParse instead groups fields into sections, associating each date field with its proper parent section.
Figure 2. Sections of the Form 1040 are shown: “Sign Here” and “Paid Preparer Use Only.” Each contains its own “Date” field.
Flat list
{ signature, "Your signature" }, { text, "Date" }, { text, "Phone no." },
{ text, "Preparer's name" }, { text, "Date" } // 54 nodes, 0 sections
Grouped into sections
{ section, "Sign Here", items: [ signature, Date, occupation, Phone ] },
{ section, "Paid Preparer Use Only", items: [ name, Date, PTIN, Phone ] }
// 28 nodes, 5 sectionsWhile checkbox state is one of the most critical pieces of information on the form, it’s a challenging task for a VLM since it is reading the entire page and predicting all of its content at once. To overcome the limitations we observe with VLMs, we built a small checkbox-state classifier, which looks at each box individually to predict the state and correct the model’s original prediction.
Figure 3 shows a ticked box that the initial VLM parse returned as false ; our classifier corrects it and reads it as checked with high confidence.
Figure 3. Checkbox-state correction on a W-9. The “Other” checkbox is marked, although the initial VLM parse records false ; the checkbox-state model corrects it to true .
before { "field": "checkbox", "label": "Other (see instructions)", "value": false }
// state model reads that box checked @ 0.94
after { "field": "checkbox", "label": "Other (see instructions)", "value": true }
// value corrected, box untouchedFinding the boxes
A bounding box is the rectangle around the object you care about, which on a form means a text field or checkbox. Bounding boxes are critical for downstream applications and verification since they let reviewers and the user check against the original document.
Figure 4. Bounding boxes on a Form 1040. Blue boxes mark text fields, and green boxes mark checkboxes.
A VLM particularly underperforms at this task and will miss lots of bounding boxes. Using a purpose-built deterministic model results in a far superior outcome. At LlamaIndex, we trained a form bounding box detection model from ground up to greatly improve performance on this task.
Figure 5 shows one observed failure: the raw VLM returned only two boxes in this run, while LlamaParse detected all field bounding boxes.
Figure 5. One observed VLM detection failure on a digital form. The VLM’s full-page output contains two misplaced boxes—one on the blank margin and one on a printed arrow (left); a crop of LlamaParse’s output on the same page shows the detected entry cells in the cost-of-goods section (right).
Even when a VLM manages to output lots of bounding boxes, aligning those detected boxes to real boxes on the page remains a challenge for VLMs.
Figure 6. Bounding-box alignment on a scanned W-4. The VLM boxes drift between entry rows (left), while LlamaParse places one box on each of the 23 entry cells (right).
There are also some open-weight models, such as FFDNet, that are built and measured on digital forms, and they can beat a VLM on accuracy. But the open models are trained and benchmarked on blank digital forms, and lack the robustness for filled data or scans. Figure 7 shows these failures on a filled W-9, while Figure 8 shows them on a scanned, hand-filled 1040.
Figure 7. Bounding-box detection on a filled W-9. The open-weight detector misses checked boxes and shifts boxes away from entered values (left); LlamaParse localizes the fields and controls on the same page (right).
Figure 8. Bounding-box detection on a scanned, hand-filled Form 1040. The open-weight detector returns one box on an empty field (left); LlamaParse detects all 51 filled fields on the same page (right).
Attributing boxes to fields
Attribution is what makes AI document processing auditable. Once bounding boxes have been extracted, they have to be linked back to the form element they correspond to.
Figure 9. The detected rectangle for line 8c is attributed to the label that governs it, “Cancellation of debt.”
A general VLM can attempt this attribution, but with limited reliability. Matching a bounding box to the right content requires understanding what each component in a form means. The complexity grows as the form becomes denser. Unfortunately, matching can fail silently even when the text is read perfectly. Figure 10 shows a VLM attribution failure: the model returns the correct value for line 8b, 70870 , but attaches its bounding box to line 8a.
Figure 10. A silent attribution error, and the same lines handled correctly. Left: a general VLM transcribes line 8b’s value correctly but attaches its bounding box to line 8a. Right: LlamaParse correctly attributes the value to line 8b.
Parsing forms with LlamaParse
Forms pose a unique set of challenges. Parsing them correctly requires a purpose-built solution that can represent the form’s structure, detect bounding boxes, and attribute each box to the right element so its source can be tracked. At LlamaIndex, we have developed a form-parsing solution that outperforms VLMs at a fraction of the price.
To try it on your own form, add one option to a parse request:
result = client.parsing.parse(
file_id=file.id,
tier="agentic",
processing_options={"forms": "enrich"},
expand=["forms"],
)
nodes = result.forms.pages[0].forms[0].json_ Enriched forms runs on the cost_effective , agentic and agentic_plus tiers and adds 10 credits per form page (it adds 0 credits for pages that don’t contain forms).
Try it out
Enriched Forms Beta is available to all customers via the API. One option turns the form pass on, and the page comes back as a JSON tree (described in more detail later):
from llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
file = client.files.create(file="w2-2024.pdf", purpose="parse")
result = client.parsing.parse(
file_id=file.id,
tier="agentic",
version="latest",
processing_options={"forms": "enrich"}, # run the form pass
expand=["forms"],
)
for page in result.forms.pages:
for form in page.forms:
for node in form.json_: # FormNode tree, in reading order
print(node)The first field of a W-2 comes back like this — the value as printed, and the box it was read from:
{
"type": "field",
"field": "text",
"id": "1",
"label": "Wages, tips, other compensation",
"value": "230303.03",
"bbox": [{"x": 356.6, "y": 143.2, "w": 96.4, "h": 14.0}]
}For more information on how to get started, try one of our cookbooks:
来源:LlamaIndex:产品、工程与评测 · llamaindex.ai