Per field, not per document
A document-level score is close to useless. If a bill scores 94%, you still do not know which of its eight fields is the shaky one, so you end up checking all of them.
Score each field separately and the picture inverts: seven values are settled and one is flagged. You look at the one.
A 76% and a 99% should not look the same
Most tools present every extracted value with identical confidence — which means every value costs the same to verify, which means the automation saved you less than it claimed.
When the difference is visible, your attention goes where the uncertainty is. That is the entire mechanism by which review gets faster.
Low scores are surfaced, never averaged away
A single uncertain field is shown to you as an uncertain field. It is not blended into a document average that makes it disappear.
The common failure mode in AI tooling is a confident-looking summary sitting on top of one bad value. Per-field scoring is the fix.
You set what needs a human
Review thresholds and rules are yours to set. What crosses the line for one firm's risk appetite will not for another's.
What is not configurable is the gate itself: every document reaches a human review queue regardless of score.