Extraction

Confidence calibration: what AWBGuru's confidence score really means

Every field AWBGuru pulls off a document carries a confidence score the model self-reports. This is what that number means, how the auto-export threshold uses it, and why "calibration" here watches real corrections instead of rescaling the score.

Sep 17, 2026· 6 min read· Confidence

scores every field it reads, so you see at a glance which values to trust and which to check — before anything exports.

Start here

When AWBGuru reads an air waybill it doesn’t just hand back values — it hands back a confidence score for every field. Each extracted field carries a number from 0 to 1 that the extraction model self-reports as it reads the document: how sure it is of that field’s value. A 1 means the model is certain; a low score means the underlying text was ambiguous, smudged, or illegible. This is field-level confidence — the model’s own per-field certainty, not a separate number from an OCR engine bolted on afterward. Understanding what that number is, and what it isn’t, is the difference between trusting straight-through export and re-checking every document by hand.

A number the model reports — not an OCR meter

The confidence score is immutable model output. As the model transcribes each field — the master air waybill number, the airport codes, the piece count, the weights — it reports how sure it is of the value it just produced. That number is stored with the field and never rewritten. It isn’t a pixel-quality reading from a scanner, and it isn’t a blended average across the document; it is one number per field, describing that field alone. A crisp, unambiguous shipper name and a barely-legible handwritten weight on the same document get their own independent scores. For how the model reads the document in the first place, see how MAWB data extraction works.

In , every field lands with its own confidence score and a color you can read without opening the record — so you know which documents can export untouched and which need a look.

Start here

The four-tier color ramp

In the UI, confidence isn’t shown as a raw decimal — it’s a four-tier color ramp on the exact rounded percentage, and the same colors follow a field everywhere it appears: capture, review, dashboard, and history. One glance tells you where to look.

how the confidence tier reads in the UI
90% and up — green, “trust”Trust
80–90% — blue, “solid”Solid
75–80% — grey, “glance”Glance
Below 75% — amber, “check it”Check it
A field the model gave no confidence for shows faint and neutral — treated as unknown, never as high.

The tiers are cut on the rounded percentage, so the ramp is consistent from screen to screen. The point isn’t the color for its own sake — it’s that a reviewer scanning the exceptions queue spots the one amber field on an otherwise-green document instantly, instead of re-reading values that were never in doubt.

Because the ramp is identical in capture, review, the dashboard, and history, a field that read amber at capture reads amber everywhere it’s shown later — the color is a property of the score, not of the screen. That consistency is what lets a team build a habit around it: green means move on, amber means look. It also means the same field-level confidence a reviewer sees in the queue is the number the auto-export gate is deciding on, with nothing translated in between.

The auto-export threshold

Colors help a human; the auto-export threshold is what lets AWBGuru skip the human entirely. It’s a single dial the workspace Owner sets, with three allowed values: 80, 85, or 90 percent. A document flows straight through — auto-exports with no review — only when every non-empty field scores at or above that threshold. If a single field falls short, the whole document routes to human review.

The rule is deliberately fail-safe. A field that is missing a confidence value, or that scored zero, counts as below the threshold — it does not slip through on a technicality. Raising the dial to 90 sends more documents to review but tightens what auto-exports; lowering it to 80 lets more through with a wider margin for a wrong value. The Owner owns that trade-off, per workspace.

Note the wording: every non-empty field. A field the document simply doesn’t carry isn’t held against the gate — it weighs the fields that were actually extracted, not the ones that were never there. What it will not do is wave through a field that should have a value but came back with none, or with zero confidence; those are exactly the cases the fail-safe is built to stop.

What “calibration” actually means here

This is where the word calibration usually gets oversold, so let’s be precise: AWBGuru does not statistically rescale the confidence score. There is no Platt scaling, no isotonic regression, no reliability-curve fitting that rewrites a 0.82 into a “true” probability. The number the model reports is the number you get. What AWBGuru calibrates is something more grounded — whether those scores are actually earning your trust in practice — and it does it by watching real human corrections.

That monitoring loop has two visible parts. The first is a retro preview when the Owner picks the threshold: had the threshold been set at a given value over the last several months, these high-confidence fields would still have needed a human fix. It shows the trade-off against the workspace’s own history before anyone commits to a number.

The second is a correction-rate circuit breaker. If a specific field gets corrected by humans in three or more distinct documents inside a rolling 30-day window — despite scoring above the threshold — that field is automatically pulled out of the auto-export path and sent to review. It’s the system noticing that a field’s high scores aren’t holding up in the real world and reacting on its own. And it’s self-healing: once those corrections age out of the 30-day window, the field recovers and returns to the auto-export path automatically. That feedback loop — watch the corrections, react, recover — is the closest thing AWBGuru has to calibration, and it’s honest about being a monitoring mechanism rather than a statistical one.

Validation runs alongside, not inside, the score

Confidence answers one question — how sure is the model of what it read. It does not check whether the values make sense together. That’s a separate, parallel layer: cross-field validation. The master air waybill number’s check digit is recomputed against the mod-7 rule; the prefix is matched to the carrier; airport and route codes are checked; chargeable-versus-gross weight is recomputed. Each finding routes the document by severity — an error rejects it, a warning sends it to review, a clean pass lets it continue — and none of it changes the numeric confidence score. The confidence number stays immutable model output; validation runs beside it.

It’s worth being exact about the shape of this, because it’s easy to picture a single blended “document score” that everything rolls up into. There isn’t one. AWBGuru runs two independent mechanisms: per-field confidence measured against the auto-export threshold, and validation findings routed by severity. A document has to satisfy both to auto-export — every field at or above the threshold, and no blocking validation error. Treating them as one number is where teams get surprised; keeping them separate is what makes “why did this go to review?” always answerable.

Put together, that’s what it means to trust auto-export in AWBGuru. A field’s confidence tells you how sure the model was; the threshold turns that into a pass/fail gate you control; the correction-monitoring loop keeps the gate honest as your real documents flow through; and validation catches the errors a confident-but-wrong value would otherwise carry downstream. No single magic number — just two clear layers you can reason about, and a feedback loop that adjusts to what your reviewers actually fix.

Frequently asked

What does the confidence score on an extracted field mean?

It’s a 0–1 number the extraction model self-reports for each field as it reads the document — how sure it is of that field’s value. A 1 means certain; a low score means the text was ambiguous or hard to read. It’s the model’s own per-field certainty, not a separate OCR-engine number.

How does a document skip human review and auto-export?

The workspace Owner sets an auto-export threshold of 80, 85, or 90 percent. A document flows straight through only when every non-empty field scores at or above it. Any field below — including one missing a confidence value or scoring zero — routes the whole document to review. It’s fail-safe, so an unscored field never slips through.

Does AWBGuru statistically recalibrate the confidence score?

No. Calibration in AWBGuru is a monitoring loop over real human corrections, not a statistical rescaling of the number. If a field scoring above the threshold gets corrected in three or more distinct documents within a 30-day window, it’s automatically pulled into review, and it recovers once those corrections age out. The confidence number the model reports is never rewritten.

See the confidence on your own documents.

Start with 30 free scans — no card. Run a few AWBs, watch every field score itself, and set the auto-export threshold you trust.

← All resources