The problem
Teams were retyping data from PDFs, scans and handwritten forms into their systems.
Context
Business documents arrived in every format: typed, scanned, photographed and handwritten. Every one was keyed in by hand, which was slow, error-prone and impossible to audit.
What I built
- An extraction pipeline that reads documents, including handwriting, and pulls out structured fields.
- Confidence thresholds: high-confidence data goes straight into the CRM or ERP, and low-confidence data goes to a person for review.
- Precedence rules for when two sources disagree, so the system always knows which value wins.
- Automatic alerts to the client’s team for any document the system can’t parse.
- Processing runs on the client’s own infrastructure, with retention rules set by the client.
Trade-off worth naming
Sending uncertain extractions to a human reviewer slows a small share of documents down, but it means nothing wrong enters the system silently.