Air waybill data extraction turns an air waybill document — a scan, a PDF, a phone photo — into a structured, validated shipment record a system can act on. The AWB carries the fields that move freight: the AWB number, origin and destination, pieces and weight, charges, and handling instructions. Extraction reads those fields and returns them as data. Whether that data is usable depends on whether it was validated against how air freight works.
Reading text off a document is the commoditized part now; plenty of tools do OCR. The harder problem with an air waybill is knowing whether what was read is correct. A confident read of the wrong AWB number is worse than a blank, because it looks trustworthy and flows straight into a booking or an invoice. OCR tells you what the document says. Air waybill data extraction tells you which fields you can safely act on.
Where the document comes from
Air waybills arrive as scans, clean PDFs, skewed phone photos, or email attachments. The first job is to normalize all of it into something the pipeline can read, over an API or by email. Forward a waybill to a dedicated address and a result comes back; in the common case there is no portal and no manual upload step. For the full path from intake to output, see what AWB data capture is, and why OCR alone isn't enough.
With , you can run air waybill data extraction on your own waybills today.
Start hereReading the fields
Extraction pulls the fields off the document: the AWB number, origin and destination, pieces, weight, nature of goods, charges. Each read carries a confidence score the model self-reports, because not every field is equally legible — a crisp printed prefix scores higher than a smudged handwritten weight. That score feeds the next step; it is not the final word on whether a field is right. For a field-by-field tour of what's on the document, see AWB data explained.
Why validation is the hard part
This is the core of air waybill data extraction — checking each read against how air freight actually behaves, not just whether the characters were legible. A few of the checks that run on every document:
- AWB check digit (MOD-7) — the serial's last digit must equal the first seven modulo 7, so a transposed or misread number is caught, not trusted. A failure is a hard reject.
- Prefix belongs to a carrier — the airline prefix must be a real IATA carrier prefix, cross-checked against the carrier and routing on the document, so a mistyped or mismatched prefix is flagged.
- Chargeable weight can't fall below gross — chargeable is the greater of actual and volumetric weight, so a chargeable below gross is physically impossible and gets flagged as a likely transpose. Wrong weight is wrong money for a forwarder.
- Routing resolves — origin and destination must be valid IATA airport or city codes, and the two can't be the same place.
- Duplicate copies are caught — a repeat AWB number is tagged as another copy of the same waybill, so one document isn't processed or billed twice.
A consolidation is also detected before anything reads the master as a single commodity. Every check resolves to one of three outcomes — clear (straight through), review (routed to a person), or reject (a hard failure like a bad check digit) — and confidence gets a reality check on top: a field can read with high confidence and still be wrong, and a field can read with low confidence and still be verified against a rule.
What comes out
What comes out of air waybill data extraction is one validated record, available as JSON or XML over the API, as FWB/16 or FWB/17 EDI, and as CSV or XLSX. Every format reads from the same record, so they agree. Because validation has already run, most documents clear on their own and only the exceptions need a person — that is what exceptions-only review means in practice. For the same four steps written out specifically for a master air waybill, see how MAWB data extraction works.