AWB data capture is the process of turning an air waybill, whether a PDF, a scan, or a phone photo from the dock, into a structured, validated shipment record that other systems can use. It covers everything from receiving the document to delivering the finished record as JSON, XML, or an FWB message, with validation and a confidence score deciding which documents need a person and which pass straight through.
That definition is wider than "OCR" on purpose. OCR reads characters. Capture has to name them, check them, score them, and hand them off. Below: what the full process covers, then a made-up waybill showing the errors that pass OCR cleanly and fail the moment freight logic is applied.
What "capture" covers
On a forwarder's desk, capture is a pipeline with six stages, and the value is in the later ones.
- Intake. An email you forward, a combined PDF you drop in (multi-page bundles are split page by page), a photo taken at the dock, or an API. The source doesn't change what happens next.
- Extraction. The page is read into around forty named fields: the AWB identity (prefix, serial, check digit), the parties (shipper, consignee, issuing agent and its IATA code), the routing (departure, destination, flights and dates for up to two legs), the cargo (pieces, gross and chargeable weight, goods, dimensions, HS code), the charge boxes, handling information, and the execution date and place.
- Validation. Deterministic air-freight rules run over those values. This is the part OCR cannot do, and the next section is about it.
- Confidence. Each field carries a score, and so does the document. You choose the threshold at which a clean, high-confidence document is released without anyone opening it.
- Exceptions-only review. A document reaches a person only when a rule fires or confidence is too low. Everything else never touches a queue. (Why that matters: exceptions-only review.)
- Outputs. The same record is rendered in whatever format the consuming system speaks: JSON, XML, CSV, XLSX, or an FWB v16 or v17 message.
With , you can run all six capture stages on your own waybills today.
Start hereWhy OCR alone isn't enough on a MAWB
Take a synthetic master waybill. Every number and name below is fabricated.
AWB 999-44719286 # 999 is a stand-in prefix, not a real carrier Shipper Northgate Precision Components, Dayton OH Consignee Harbor Ridge Distribution, Yokohama Routing ORD → NRT flight XX 412 / 14 Sep Pieces 18 Gross wt 312.0 K Chg wt 213.0 Rate class Q Currency USD
An OCR engine reads every one of those characters at high confidence, and it should: the print is crisp. Now look at what a validator sees.
1. The check digit doesn't match the serial
An AWB number is a 3-digit prefix, a 7-digit serial, and one check digit that must equal the serial modulo 7. For 4471928 the remainder is 6, and the printed digit is 6, so the number is consistent. Transpose two digits during a re-key, 4417928, and the remainder becomes 4 while the page still says 6. Both versions are legible digits; only arithmetic notices. A failed check digit is an error, not a warning: the document cannot pass, and it lands in the exception queue with the number flagged, because that number would mis-route everything downstream. (Full mechanics: how the MOD-7 rule works.)
2. The airline prefix disagrees with the flight
The three-digit prefix names an airline. A prefix that isn't three digits is structurally impossible and is flagged as an error. One that is well-formed but not in the airline table is flagged as a warning, because free reference tables can lag and a false unknown is worse than no check. The prefix is also cross-checked against the two-letter carrier code at the front of the flight number; carriers hold sets of prefixes, so it's a membership test. If the prefix belongs to one carrier and the flight says another, the document goes to review. OCR has no opinion on any of this. It read "999" and "XX 412" clearly, and that was the whole job.
3. The routing points somewhere that doesn't exist
Departure and destination are checked against the IATA airport and city table. A zero misread as the letter O turns ORD into a code that isn't an airport, and OCR confidence won't flag it, because an O is a perfectly good character. The validator also catches a departure equal to the destination, and a routing that is simply missing. Each routes the document to a person instead of into an EDI message a carrier system will bounce.
4. The weights are physically impossible
Chargeable weight is the greater of actual and volumetric weight, so it can never be below gross. On the sample, gross is 312.0 and chargeable is 213.0: a transposition, and one that costs money, because the rate is applied to the chargeable figure. The rule fires only when both values are present and the relationship is impossible, so a blank chargeable box produces no noise. (Why the two weights differ at all: chargeable vs gross weight.)
5. The charge boxes say something the codes don't allow
Format checks cover the charge declarations: currency must be an ISO 4217 code, rate class must be one of the IATA rate class letters, and an HS code must have 6, 8, or 10 digits. Just as important is what is deliberately not checked. Blank charge boxes are never flagged, because a consolidation master legitimately leaves them empty, and a rules engine that sends every blank Total to review buries the queue in false exceptions within a week.
The pattern across all five: a valid-looking value in the wrong place, or two valid-looking values that contradict each other. OCR confidence measures legibility. Validation measures whether the shipment makes sense. A capture system needs both, and only the second decides whether a person is needed: an error or a warning routes the document to review, and a document with neither passes.
Forwarder-side "AWB capture" vs airline-side "AWB data capture"
The phrase is used on both ends of the same handoff. On the airline side, "AWB data capture" usually means the carrier's cargo system receiving, or a cargo agent keying, the FWB message into the carrier's own records. On the forwarder side, "AWB capture" means starting from the document you hold and producing a record you can bill, file, and transmit from. AWBGuru sits on the forwarder side: it produces the record and generates the FWB message; transmission to the carrier stays on whatever channel you already use.
What good capture output looks like
The test of a capture system is that one record feeds every format without re-keying: a JSON body for your API integration, an XML or CSV file for a system that wants a file, and a positional FWB message for the carrier side, each rendered from the one record rather than re-read from the page. An illustrative slice, using the field names from the sample above:
{
"AWB2": "999-44719286",
"Departure": "ORD",
"To": "NRT",
"Pcs": "18",
"Weight": "312.0K",
"Charge_Weight": "213.0",
"Currency": "USD",
"outcome": "Review" # chargeable weight below gross
}
Two things make this usable rather than merely complete. The outcome travels with the data, so a consuming system knows whether a human has confirmed it. And the field names are the same on every document, whichever airline's form it was printed on, so there is nothing to map and nothing to re-map when a carrier changes its layout. That is the difference between a capture record and an OCR text dump. (Full format list: outputs.)