Legal teams do not have a document-processing problem. They have a liability problem. An extraction that is 95% right and cannot tell you which 5% it got wrong is not a time saving. It is a new review task.
The pattern we deploy
- Layout-aware parsing first: clause boundaries, defined terms and cross-references survive the trip out of the PDF.
- Every extracted field carries a span citation back to the page and line it came from. One click jumps there.
- Calibrated confidence, not model logprobs. Anything below threshold routes to review rather than being asserted.
- The reviewer's corrections become the eval set, so the threshold moves in the right direction each quarter.
Where it pays off
Diligence rooms, renewal and obligation tracking, and any portfolio where the same twelve clauses decide most of the risk. It pays off worst on bespoke, heavily-negotiated one-offs. Those are still a lawyer's job, and should be.