AI extraction of supplier emails with a human review queue
Scenario: distributor receiving orders, invoices, and price updates by email
A worked design for the most requested AI use case in operations: turning the emails your team retypes every morning into structured, verified data — without ever letting a language model write unchecked numbers into your books.
The problem
Distributors and wholesalers receive a steady stream of supplier emails: order confirmations, invoices as PDFs, price-list updates, delivery changes. Someone reads each one and retypes the contents into the ERP or accounting system. It's hours of daily work, it's error-prone in exactly the places errors are expensive (quantities, prices, references), and it doesn't scale with the business. Generic 'AI email assistants' fail here for one reason: they're confidently wrong just often enough that nobody trusts them with the books.
Requirements
- Extract structured data (supplier, references, line items, amounts, dates) from emails and PDF attachments
- No value reaches a business system unless it passed validation and confidence checks — or a human confirmed it
- Uncertain extractions become a review task, not a silent guess
- Every written record links back to the exact source email for audit
- Running cost visible per document type, so AI spend is a number, not a fear
The solution
A pipeline, not a chatbot: emails are normalized (attachments parsed, noise stripped), then an LLM extracts against a strict typed schema. Schema validation and confidence scoring decide the path — high-confidence documents write to the ERP/accounting system through the same idempotent machinery as our sync designs; anything uncertain lands in a review queue where a human confirms once, and the correction feeds back as an example. The model proposes; the schema and the human dispose.
Architecture
Incoming mail hits a normalization step (MIME parsing, PDF text extraction, supplier identification), then the extraction service calls the LLM with a typed schema per document type. Results pass through declared validation rules (totals must sum, references must match known formats, dates must parse) plus a confidence threshold. Passing documents are written to business systems via idempotent workers; failing or low-confidence ones enter the review queue with the source displayed alongside the extraction. Every step — prompt, model output, decision, human correction — is recorded in an audit log, and a cost monitor aggregates spend per supplier and document type.
Challenges & resolutions
Language models are occasionally confidently wrong — and one silently wrong invoice amount costs more than a month of manual retyping saved.
Resolution — Nothing the model outputs is trusted directly. Extractions must pass declared validation rules (line items sum to the total, references match the supplier's known format) and a confidence threshold; anything else goes to the review queue. The failure mode is a queued task for a human, never a wrong number in the books.
Per-document API costs look tiny in a demo and compound alarmingly at volume.
Resolution — Cost is designed in, not discovered later: documents are routed to the smallest model that handles their type, repeated formats hit cached extraction patterns, and a monitor reports spend per supplier and document type — so the business sees 'invoices cost $0.0x each' as a fact on a dashboard.
Supplier formats vary wildly — pristine PDFs, photographed delivery notes, prose emails with the order buried in paragraph three.
Resolution — The normalization step does the unglamorous work before any model sees the document: attachment extraction, OCR where needed, supplier identification. Document types with hopeless quality are configured to route straight to the review queue — the design degrades to 'organized manual entry', never to guessing.
Why this design exists
"Someone spends every morning retyping supplier emails" is the most common sentence in conversations about AI and operations. The technology genuinely fits — but only inside an engineering frame that treats the model as a fallible component, not an oracle. This worked design is that frame, spelled out end to end.
Design notes
The model proposes; the schema disposes. The LLM's job is to fill a strict typed schema, nothing more. Validation rules — totals that must sum, reference formats that must match, dates that must parse — are ordinary deterministic code. The intelligence is bounded by the same discipline as any other input: never trust, always verify.
The review queue is a feature, not an apology. A percentage of documents will always be ambiguous. Designing the human step in from the start — source shown beside extraction, one-click confirm or correct — turns "AI that can't be trusted" into "AI that does the typing while a person does the judging."
Idempotent writes, as always. The write side reuses the queue-and-worker skeleton from our Shopify–ERP sync design: idempotency keys, retries with backoff, an audit trail. An extraction pipeline that can double-post an invoice on retry hasn't removed manual work — it has created reconciliation work.
Provider-agnostic on purpose. The extraction interface speaks to Claude by default and treats the provider as an adapter — models improve and prices move, and the design shouldn't need a rebuild to follow them.
What we'd adapt per client
The document types and their schemas (orders, invoices, price lists, delivery notes), the confidence thresholds per type — set conservatively at first and loosened as review-queue history proves accuracy — and the target systems on the write side. The pipeline shape, the review queue, and the cost monitoring carry over unchanged.
What the design guarantees
- No value reaches a business system without passing validation and confidence checks, or explicit human confirmation — by construction, not by policy
- Every written record links to its source email and its extraction, so any number can be audited in seconds
- Uncertain documents cost one human confirmation instead of one silent error
- AI spend is a visible per-document number, with model routing keeping it proportionate
Have a similar problem?
Describe the problem in plain language — broken, slow, manual, or missing. We'll tell you honestly whether and how we can help.
- 01We reply within one business day
- 02A short call to understand the problem
- 03A written scope and fixed quote — no obligation