Documents in, validated data out. People decide the exceptions.
Applications, invoices, claims and letters arrive in every format. We build intake flows that classify each document, extract the fields that matter, check them against your rules and send only the uncertain cases to a person.
For organisations that receive documents in several languages, formats and countries, and need every value to be traceable.
Documents are captured from every channel, classified, extracted and validated by the system. Cases below the confidence threshold or failing a rule go to a person in a review queue. Validated data is handed off to the business system, and reviewer corrections feed the evaluation set.
Before and after
What changes in the working day
Today
With the system
Today
Today: Retyping from attachments
Staff open each PDF, scan or e-form and copy names, amounts and references into the case or ERP system by hand.
With the system
With the system: Data arrives structured
Fields are extracted with their position in the source document, so a reviewer sees the value and the evidence side by side.
Today
Today: Errors found late
Missing pages and inconsistent data surface weeks later, when a decision is challenged or a payment bounces.
With the system
With the system: Rules run at intake
Completeness, master-data matches and business rules are checked before a case is opened, while the sender can still be asked.
Today
Today: Every document gets the same attention
A clean, complete invoice takes as long to process as a messy, ambiguous one.
With the system
With the system: Attention goes to the exceptions
Clean documents pass straight through. People work a queue of the uncertain and unusual cases, with the reason each one is there.
Today
Today: No record of how data was captured
When an auditor asks where a value came from, the answer depends on who keyed it in.
With the system
With the system: Every value is traceable
Source document, extraction, confidence, rule results and reviewer changes are logged for each field.
Workflow
How the workflow runs
Seven steps, one exception path. The thresholds and rules are agreed with your process owners and tuned on your own documents before anything runs in production.
Illustrative workflow
7 steps · 1 human checkpoint
Legend
System
Person
Human checkpoint
Exception path
Document intelligence: illustrative workflow with exception path
Read the diagram as text
The main path runs from capture through classification, extraction and validation to the hand-off into the business system, and ends with learning from corrections.
After validation, any document with a field below the confidence threshold or a failed business rule leaves the main path and goes to the review queue, where a person confirms or corrects the values. Reviewed documents then rejoin the main path at the hand-off step.
The review queue is the human checkpoint: uncertain values never reach the business system without a person.
01
Step 1: Capture from every channelSystem
Scans, email attachments, portal uploads and e-forms land in one intake queue with a case reference, channel metadata and a checksum. Structured e-invoices (Peppol UBL) bypass extraction and go straight to validation.
02
Step 2: Split and classifySystem
Multi-document bundles are split into their parts. Each part is classified by document type and language, which selects the extraction schema and the rule set that apply.
03
Step 3: Extract fieldsSystem
Layout-aware OCR reads the page; a language model extracts the agreed fields into a strict schema. Every field carries a confidence score and its location in the source.
04
Step 4: Validate against rulesSystem
Values are checked against master data (customer, supplier, VAT and IBAN registers), cross-field rules and duplicates. Fields above the agreed confidence threshold that pass every rule go straight through.
05
Step 5: Review queuePerson
A reviewer sees the document and the extracted values side by side, with the failed rule or low-confidence field highlighted, and confirms or corrects them.
Exception path: Below threshold or rule failed· Returns to step 6
Human checkpoint
Uncertain values never reach the business system without a person. Reviewers can also reject a document and request a new one from the sender.
06
Step 6: Hand off to the business systemSystem
Validated data opens or updates the case in the case-management system, CRM or ERP through its API, with the source document attached and a retention class set.
07
Step 7: Learn from correctionsSystem
Reviewer corrections are captured as labelled examples. They extend the evaluation set, and thresholds or prompts change only after the new version passes it.
Components
What we would build
01
Intake and channel connectors
Mailbox, scanner, portal and e-form connectors that deliver every document to one queue with metadata and a case reference.
Reviewers only see the document types and cases their role allows. Extraction services run without standing access to the business system; the hand-off uses a scoped service account per integration.
Human review
Confidence thresholds and rule failures route cases to a person. Thresholds are set per field, not per document, and are only lowered after the evaluation set shows it is safe. Decisions on applications stay with case handlers.
Auditability
For every field we log the source location, extracted value, confidence, rule results, model and prompt version, and any reviewer change with reason code.
Data protection
Personal data is minimised at intake, redaction is available for disclosure requests, and retention and deletion follow your records schedule. Processing stays in the agreed region.
AI transparency
Model and prompt versions are documented with their evaluation results. Where outputs affect people, the use of AI is recorded in your AI register and explained in the notice to applicants.
Measures
What we would measure
We agree these measures with you during discovery and record a baseline from your current process first.
What we would measure
Metric
Why it matters
How we would measure it
01Straight-through rate
Why it mattersShows how much routine work no longer needs a person, per document type.
How we would measure itShare of documents that pass every threshold and rule without review, from workflow logs.
02Exception rate and reasons
Why it mattersTells you whether exceptions come from bad scans, missing data or rules that need changing.
How we would measure itReview-queue volume by reason code, per channel and document type.
03Field-level accuracy
Why it mattersThe quality measure that matters for downstream decisions and payments.
How we would measure itPrecision and recall per field on an agreed, labelled validation set, re-run at every release.
04Handling time
Why it mattersCaptures the effort saved and the effort moved into review.
How we would measure itTime from receipt to hand-off, and active review time per exception, from queue timestamps.
No targets are set before a baseline exists.
Rollout
How we would roll it out
Phase 01
Discovery and baseline
Two or three document types, one channel, your current volumes and error patterns. We label a validation set from real, anonymised documents.
Exit criteria
Fields, rules and quality criteria agreed
Baseline measured
Data protection impact assessment started
Phase 02
Pilot in shadow mode
The system extracts and validates in parallel with the current process; people keep doing the work and compare.
Exit criteria
Field accuracy meets the agreed criteria
Thresholds set per field
Reviewers trained on the workbench
Phase 03
Controlled production
Straight-through processing switched on for the piloted document types, with sampling of auto-approved cases.
Exit criteria
Sampling shows no unacceptable errors
Runbook and monitoring handed over
Rollback tested
Phase 04
Scale
More document types, channels and languages, each passing the same evaluation gate before it goes live.
Delivery notes, certificates of conformity and customs documents matched to purchase orders.
Illustrative example
Supplier invoices across three subsidiaries
Situation
A group with subsidiaries in three EU countries receives supplier invoices as PDFs, scans and structured e-invoices, in three languages, into separate finance mailboxes.
System
One intake routes structured e-invoices straight to validation and extracts the rest. Values are matched to purchase orders and supplier master data in the group ERP.
Human control
Invoices with low-confidence fields, unknown bank details or a price mismatch wait in a review queue for the subsidiary finance team. Payments are released only from the ERP.
What we would measure
Straight-through rate per subsidiary and channel, exception reasons, and field accuracy on a labelled validation set.
International
Same invoice, three formats
Organisations operating in several EU Member States receive the same document type in different languages, layouts and legal formats. We design one intake and extraction layer with schemas and rule sets per country, so a Belgian invoice, a German delivery note and a French form follow the same controlled path.
E-invoicing mandates are converging on structured formats under the EU VAT in the Digital Age package, so structured invoices skip extraction altogether. Where extracted data feeds decisions about people, we design with GDPR and the EU AI Act in mind: documented processing, human review and logs your auditors can follow.
It depends on your documents, so we measure it rather than promise it. We label a validation set from your own anonymised documents and report precision and recall per field. Extraction models make mistakes, which is why uncertain fields go to a person instead of straight into the business system.
02Can it handle handwriting, poor scans and several languages?
Handwritten forms and poor scans can be read, with lower confidence, so more of them land in the review queue. Dutch, French, German and English documents are classified by language first, so each one gets the right schema and rules.
03We already have OCR or a capture product. Do we replace it?
Not necessarily. If your capture layer works, we build validation, review and hand-off around it, and add model-based extraction only for the document types where it adds value.
04Does the system make decisions on applications or claims?
No. It prepares and checks data. Decisions with legal or financial effect stay with your case handlers, in line with GDPR article 22 and your own policies.
05Where are documents processed and stored?
In the cloud region or data centre you choose. We document every processing step, including any external model service, so your data protection officer can assess it before go-live.
06What is a sensible first scope?
One high-volume document type with clear rules, such as supplier invoices or a single application form. It proves the workflow, the review screen and the measurements before more types are added. Read about our delivery approach.
Discuss your document flows
Bring two or three document types and the systems they end up in. We will show what a first release could look like across your markets and languages.