抽取置信度分数实用指南:如何设定 cutoff 平衡精度与人工审核
A practical guide to extraction confidence scores
LlamaIndex 在 Extract 中为抽取结果提供 0 到 1 的置信度分数,用户可设定 cutoff 决定哪些值自动通过、哪些转人工审核。在 ExtractBench 的 370 份文档上,Extract Agentic Plus 以约 0.77 的 cutoff 达到 97% 精度,自动接受约 66% 的期望字段。
Extracting data from documents at scale comes with the risk of incorrect or hallucinated values. In many enterprise workflows, human-in-the-loop review is a necessary step in the process. That raises a practical question: how do you draw a line between values you can reliably pass to the next step in your workflow and those that need closer review?
In Extract, we offer confidence scores alongside your extraction results. Each score ranges from 0 to 1. A higher score indicates greater confidence that the extracted value is correct.
You define the line between fully automated results and those needing review using a confidence score cutoff. Values with confidence scores at or above the cutoff are accepted automatically; values with lower scores go to review.
Choose your accuracy target
Suppose your workflow needs at least 97% of automatically accepted values to be correct. Here’s how to find a confidence cutoff that meets that target while keeping manual review to a minimum.
- Create a representative dataset of the documents you will be extracting from in production.
- Define the ground truth: the correct values for all expected fields in that dataset.
- Run extraction on all the documents in the dataset to get extraction results and confidence scores.
- Start with a low confidence score cutoff and increase it until the accepted values meet your precision target.
- Check that the cutoff still meets your precision target on another representative dataset that wasn't used to choose it.
- Periodically rerun this process as the profile of your production documents changes to check that your cutoff still meets your precision target.
Finding the lowest cutoff that meets your precision target matters because it lets you send more values through automatically without review.
Measure how much you can automate
Let’s look at Extract’s results on ExtractBench. At 97% precision, how much data can we accept automatically?
For Extract Agentic Plus, a confidence cutoff of about 0.77 achieved 97% precision. Move the cutoff below to see how the accepted values and review workload change.
PortableText [components.type] is missing "confidenceFigure"
Two measures describe the tradeoff at this cutoff:
Precision after filtering: the share of accepted values that are correct.
559,666⏞Accepted correctly576,975⏟All accepted values≈97.0%
Recall after filtering: the share of all expected fields accepted correctly.
559,666⏞Accepted correctly841,895⏟All expected fields≈66.48%
At this cutoff, roughly two-thirds of expected fields were accepted correctly without review, at 97% precision.
Let’s say invoice totals need 97% precision, while product descriptions are fine with 90%. Choose these targets based on what an error would cost and how much review your team can handle. You can use this same process to choose separate cutoffs for fields or groups of fields in your schema.
What makes a confidence score useful?
To be useful, confidence scores should be higher for correct values and lower for likely errors. That lets you keep more correct values above the cutoff while sending uncertain ones for review.
That benefit also depends on coverage: how many returned values receive a score. Unscored values need review or a separate acceptance rule, so missing scores put a ceiling on what confidence filtering alone can automate.
How smoothly you can adjust that workload depends on granularity. Values with the same score cross the cutoff together, so a large tied group can make a small cutoff change produce a sharp jump in recall and review volume. More distinct scores let you make smaller adjustments.
Comparing confidence scores on ExtractBench
Alongside extraction accuracy, coverage and granularity determine how much correct data a system can accept at your required precision. We compared Extract, Reducto Deep Extract, and Extend with Review Agent on the same 370 ExtractBench documents.
For each system, we selected the lowest cutoff that met each precision target on the pooled field results. The chart shows how much correct data the cutoff automatically accepted.
PortableText [components.type] is missing "confidenceFigure"
At a 97% precision target, Agentic Plus accepts 66% of expected fields correctly, versus 33% for Reducto. Extend has no qualifying cutoff. These benchmark cutoffs were selected retrospectively; choose and verify yours on separate representative samples.
The score distributions below help explain the differences. Agentic Plus supplied scores for every returned field and had 770,431 distinct scores, allowing finer adjustments than Extend's five score levels.
PortableText [components.type] is missing "confidenceFigure"
The precision-target chart compares individual accuracy targets. To compare performance across all acceptance levels, we used AUGRC, the Area Under the Generalized Risk Coverage Curve. It summarizes the risk of errors passing through without review into one number. Lower is better. See how AUGRC is calculated.
Agentic Plus had an AUGRC × 1,000 of 15.0, the lowest in this comparison. A random acceptance order of the same results gives 26.5.
The dot plot pairs this risk measure with billed cost per page, so you can weigh extraction quality against what it costs to run.
PortableText [components.type] is missing "confidenceFigure"
The Agentic tier also achieved a lower AUGRC at a lower billed cost per page than either competitor.
Extract sets the standard for usable extraction confidence scores. In our evaluation, Agentic Plus automatically accepted substantially more correct data than the other APIs across most precision targets.
Try confidence scores in Extract on your own documents to see how much you can automate at the accuracy your system requires.
Get started with Extract
Run the Python example to extract a sample invoice with confidence scores and citations.
PortableText [components.type] is missing "confidenceDetails"
PortableText [components.type] is missing "confidenceDetails"
PortableText [components.type] is missing "confidenceDetails"
来源:LlamaIndex:产品、工程与评测 · llamaindex.ai