Compare AWB capture software on two things: how completely it reads a master air waybill, and what it does with the values once it has them. Reading is close to table stakes across the category now; the real spread is in validation, exception routing, output formats, pricing, and data handling. This guide sorts the category into three kinds, gives you eight criteria to judge them on, and is honest about when each kind is the right choice.
What AWB capture software is, and the three kinds you'll meet
AWB capture software (sometimes sold as AWB data capture or AWB OCR) turns a scanned, photographed, or emailed air waybill into structured data: the AWB number, parties, routing, pieces and weights, the rate line and charges, and the handling text, delivered to a TMS, a spreadsheet, or an EDI message instead of being re-keyed. Products that do this fall into three groups.
1. General-purpose document-AI platforms
Vendors such as V7 Labs, Klippa, Nanonets, Rossum, and Veryfi market intelligent document processing platforms built for many document types: invoices, receipts, contracts, forms, and in several cases logistics paperwork including air waybills. The design center is breadth: you pick or define a document type, extract the fields, and integrate the result into whatever workflow you already run. Plenty of forwarders use them well when air waybills are one document among many.
2. Logistics-document OCR vendors
A second group narrows the scope to freight paperwork: bills of lading, commercial invoices, packing lists, air waybills. Field sets sit closer to a forwarder's needs out of the box. The output is typically still the extracted fields plus a confidence score, handed back for your own systems to check.
3. A freight-native validating engine
The third category, where AWBGuru sits, treats extraction as the first half of the job. After the read, each field is checked against the rules air cargo runs on: the MOD-7 check digit, prefix, airport codes, weights, currency and rate-class codes. Each check resolves to pass, warn, or fail, and the outcome decides whether the document clears on its own or routes to review. The difference is not OCR quality. It is extract-and-hand-back versus extract-validate-and-route.
With , you can see what a validating engine does to your own waybills, not just read them.
Start hereA comparison framework: eight criteria
Eight questions a forwarder should put to any candidate, roughly in the order they matter once the product is live.
Walking the criteria
Field coverage
General-purpose platforms typically let you define your own fields, which is flexible but means someone on your side must know the IATA form well enough to specify all of them. AWBGuru's field catalog is fixed and freight-specific: 40 fields spanning the prefix and serial, shipper, consignee, and agent blocks, the agent's IATA code, departure and first routing point, both flight legs, nature of goods verbatim, rate class, HS code, dimensions, pieces and weights, the rate line and charge summary, declared values, handling information, and the executed date and place. Nothing to configure; the same catalog applies to every document. The read itself is described in how MAWB data extraction works.
Freight-rule validation
This is where the categories part ways. OCR-first tools typically return the value and a confidence score; any rule you want applied lives in your own code. A validating engine applies the rules before the data leaves. AWBGuru's checks include: the AWB serial must be eight digits and its last digit must equal the first seven modulo 7 (the MOD-7 check digit); the airline prefix must be exactly three digits and is looked up in an IATA airline table; the carrier code at the front of the flight number must own that prefix; departure and first routing point must be known IATA locations and must not be the same place; chargeable weight cannot be below gross weight (chargeable vs gross); currency must be an ISO 4217 code; an HS code must have 6, 8, or 10 digits; pieces must be a positive integer.
Severity is deliberate: a structurally impossible value, such as a failing check digit, is an error; an unknown-in-table value, such as a prefix the airline table does not list, is only a warning, because reference tables lag and a false rejection is worse than a soft flag. Blank charge fields are never flagged; they are legitimately blank on consolidation masters and plenty of ordinary AWBs.
Confidence and exceptions handling
Per-field confidence is common across the category; what differs is what happens with it. In AWBGuru, a non-empty field below 60% confidence raises a warning, and two or more fields at or below 75% raise another, since one shaky field can ride on its neighbors but two makes the whole read suspect. A warning routes the document to review; an error rejects it; a document with neither passes straight through. Why a confidence score alone is not enough is the subject of validation that catches what OCR confidence misses; the review model is described in exceptions-only review.
Output formats
General-purpose platforms typically return JSON, sometimes CSV, with EDI left to you. AWBGuru renders one normalized record as JSON or XML over the API, as a Cargo-IMP FWB/16 or FWB/17 message, or as CSV or XLSX, so the formats cannot disagree. The FWB message is generated and returned; you transmit it over your own rails, and AWBGuru holds no airline credentials. Full list on outputs.
Intake paths
Most tools accept an upload and an API call. AWBGuru adds a per-account email intake address, a bulk PDF drop that splits a stacked file into one record per waybill, phone capture that crops and deskews on the device, and an API POST with a webhook on completion. Every path lands in the same validation; details on document intake.
Templates and training
Some general-purpose platforms are built around training a model on your own samples, which suits teams with unusual documents and time to label them. AWBGuru has no per-carrier template configuration and no model for you to train; the same extraction and rules apply to a clean PDF, a skewed phone photo, or a handwritten field.
Pricing model
Pricing across the category varies widely: per page, per document, per seat, platform minimums, or a mix. AWBGuru bills per scan, where one scan is one extracted master air waybill: 30 free scans to start with no card, then pay-as-you-go or a monthly plan, and no per-seat charge. House waybills, manifests, and invoices are auto-classified but not extracted and count as a fraction of a scan. Current rates on pricing.
Data handling
Ask where documents are hosted, whether they train models, and what deletion looks like. AWBGuru is hosted in the United States and offered to US-based businesses. Extracted data and source files are kept as your system of record until you delete them or close the account; on written request or closure, customer content is deleted from production systems within 30 days. De-identified data may be used to improve AWBGuru's own extraction, on by default with an opt-out, and is retained for up to 24 months.
When an OCR-first tool is the right choice
- Mixed document types dominate. If invoices, packing lists, and customs forms are the bulk of the stream and an air waybill is occasional, one platform for everything beats a specialist for one form.
- You already own a rules engine. If your TMS validates check digits, prefixes, and routing on entry, you want raw fields fast, not a second validator.
- Most volume is not air cargo. Ocean bills of lading and trucking paperwork have their own rules; an engine built for the IATA form adds little there.
- Custom training is a feature for you. Unusual internal forms, and a team with samples and time to label them, get real value from a trainable platform.
When a validating engine is the right choice
- Air waybills are the daily volume and errors are expensive downstream: a rejected EDI message, a mis-routed shipment, a billing dispute over chargeable weight.
- You do not want to write or maintain validation code. The rules are published; they should arrive already applied.
- The output has to be an FWB message or a spreadsheet, not JSON alone, and it has to agree with the API response.
- The team is small. An exceptions-only queue lets attention scale with the number of problems, not documents.
Neither answer is wrong. The mistake is buying on OCR accuracy alone and discovering the validation gap in production.