Validation

Validation that catches what OCR confidence misses

A confidence score tells you the model is sure it read the field. It says nothing about whether the value is right. Here's the second layer that closes that gap.

Aug 27, 2026· 4 min read· Validation

turns air waybills into clean, validated shipment records — no keying, no OCR clean-up.

Start here

An OCR confidence score answers one question: how sure is the model that it read those characters correctly? It does not tell you whether the value it read is the right value. Those are two different problems, and on an air waybill the second one is where the expensive mistakes live — the kind that clear your pipeline looking clean and resurface days later as a customs hold, a re-tender, or a billing dispute.

A number can be printed cleanly, scanned cleanly, and read at 99% confidence — and still be wrong. Someone keyed a transposed serial into the shipping system. An origin airport was mislabeled at tender. A commodity code was truncated. The OCR did its job perfectly: it faithfully reproduced a value that was already incorrect on the page. Confidence stays high because the model isn't guessing about the glyphs — it just has no way to know the glyphs spell a mistake.

The gap OCR confidence leaves open

Extraction confidence is a per-character or per-field probability that the characters match the pixels. It's genuinely useful — it flags smudged, low-resolution, or ambiguous glyphs so you can route those to a human. But it is blind to a whole class of errors: the ones that are perfectly legible and simply wrong. A crisp 998 where a real prefix would be 999 reads at full confidence. There is nothing blurry about it.

AWBGuru adds a second layer on top of extraction: freight-native validation. Instead of asking "did we read this correctly?", each check asks "is this value possible — does it obey the rules air cargo actually runs on?" These aren't a spell-check. They're arithmetic and reference lookups that a wrong value physically cannot satisfy.

  • MOD-7 AWB check digit — the serial's last digit must equal the 7-digit serial modulo 7.
  • IATA airport / route cross-check — origin and destination must be valid codes and match the printed city.
  • HS-code resolution — the commodity code must resolve to a real Harmonized System reference.
  • Volumetric weight recompute — chargeable weight is recomputed from the dimensions and reconciled against the stated figure.
  • Consolidation detection — is this MAWB a consolidation of house shipments, decided before single-commodity logic runs?

None of these describes an IATA certification or endorsement — they describe the published air-cargo standards themselves, applied as checks. That's the point: the rules exist independently of any one value on the page, so they can catch a value the OCR was perfectly happy with.

With , you can validate what OCR confidence misses, routing only real exceptions to a person.

Start here

Example 1 — a valid-looking AWB number that fails the arithmetic

Take the synthetic master air waybill 999-12345675, ICN → FRA. The prefix 999 is a synthetic stand-in — not a real carrier — and the number is eleven clean digits. OCR reads it at high confidence — every character is unambiguous. But the AWB number carries its own proof inside it. The last digit is a check digit: it must equal the 7-digit serial taken modulo 7.

mod-7 check · 999-12345675
# serial = 1234567, printed check digit = 5
1234567 mod 7 = 5   ✓ matches printed check digit → PASS

# now a transposed serial the OCR reads just as cleanly:
1234576 mod 7 = 0   ✗ printed digit is 5 → FAIL

The transposition (…567 → …576) is invisible to a confidence score — both versions are crisp digits. It is not invisible to the arithmetic. The check digit no longer agrees with the serial, so the field fails outright and is held before it can reach your systems — before a mis-keyed AWB number becomes a mis-routed shipment or a rejected EDI message downstream.

Example 2 — a real airport code that doesn't match the printed city

Consider a waybill where the origin field reads OSL but the printed origin city says SEOUL. OSL is a perfectly valid IATA code — it's Oslo — so a naive "is this a real airport?" lookup passes it, and OCR read the three letters cleanly. The value is legible, valid, and wrong.

route cross-check · printed city vs code
# field: origin = OSL, printed origin city = "SEOUL"
OSL → Oslo Gardermoen, serves OSLO   (valid IATA code)
printed city = SEOUL                 → Seoul metro (SEL): ICN, GMP
OSLO ≠ SEOUL metro                   ✗ does not serve city → REVIEW

# the correct reading:
ICN → Incheon Intl, serves SEOUL     (metro SEL)
printed city = SEOUL                 ✓ serves printed city → PASS

Note what "matches" means here: the airport serves the printed city, not that the two names are identical. Incheon International sits in Incheon, a separate city outside Seoul proper, but IATA files it under the Seoul metropolitan code SEL alongside Gimpo (GMP) — so a waybill printed SEOUL with ICN is correct, and one printed SEOUL with OSL is not.

This is the reconciliation a plain code lookup skips. Validating the code against the printed city name is what turns a "valid but wrong" airport into a caught exception instead of a silently delivered error — the kind that only reveals itself when the freight lands somewhere it shouldn't.

How each check feeds a confidence score

Every check resolves to one of three states — pass (the rule is satisfied), warn (plausible but unproven), or fail (a rule is broken outright) — and those results roll up into a single document confidence score. This score is freight-aware: it reflects whether the values obey air-cargo rules, not merely whether the characters were legible.

You set one threshold. Documents that clear it deliver automatically; any document with a field that warns below your cutoff — or that fails a hard rule like the check digit — routes to a review queue instead. Raise the threshold to send more to review for tighter control; lower it to auto-clear more for throughput. It's one dial, not a rules engine to maintain.

The result: the errors an OCR score is structurally unable to see get caught before they cost you. Extraction reads the fields; validation proves them — so a clean waybill clears itself, and your team spends its attention on the genuinely hard ones instead of re-checking the routine.

Frequently asked

Does a high OCR confidence score mean the extracted data is correct?

No. OCR confidence measures how sure the model is that it read the printed characters — not whether the value is right. A value can be read at high confidence and still be a mis-keyed serial or a valid-but-wrong airport code. Validation is a separate layer that checks the value against air-cargo rules.

What is the MOD-7 air waybill check digit?

The last digit of an AWB serial is a check digit equal to the 7-digit serial modulo 7. Recomputing it is pure arithmetic, so a transposed or mis-keyed serial fails the check regardless of how cleanly the characters were read.

How does the confidence score decide what a human reviews?

Each check resolves to pass, warn, or fail, and those roll up into one document confidence score. You set a single threshold: documents above it auto-clear, and any with a field below it — or a hard failure — route to review instead of delivering.

See the checks run on your own waybills.

Start with 30 free scans — no card. Watch a caught error route to review while the clean ones clear.

← All resources