A specialty insurance firm underwriting commercial property policies received submission packets as PDF documents. Each packet contained an ACORD application, loss runs from prior carriers, a statement of values, building diagrams, and supplemental questionnaires. A typical packet was forty to one hundred and twenty pages. The firm received approximately eight hundred new submissions per week.
Underwriters spent sixty percent of their time extracting data from these packets. They read the documents, identified the relevant fields (occupancy type, construction class, square footage, prior loss history, coverage limits) and entered the data into the underwriting system manually. The extraction was tedious, error-prone, and slow. A complex submission could take ninety minutes to extract. Simple submissions took thirty minutes. The firm employed fourteen underwriters, and eight of them were effectively data entry specialists for more than half their working hours.
The firm had tried optical character recognition twice. The first attempt used a general-purpose OCR tool that converted PDF pages to text. The text output was unstructured: a wall of words with no field identification. An underwriter could read the original PDF faster than parsing the OCR output. The second attempt used a template-based extraction tool that required defining field locations for each document type. The templates broke whenever a broker formatted a document differently, which happened with approximately forty percent of submissions because the firm worked with over two hundred brokers, each with their own document templates.
Both attempts failed because they treated extraction as a text conversion problem. The challenge was not converting PDF to text. The challenge was understanding what the text meant in the context of insurance underwriting.
The Extraction Problem, Reframed
We reframed the problem. The underwriter did not need the text from the document. The underwriter needed specific data points in a specific structure that the underwriting system could consume. The document was a container. The data points were the contents. Extraction meant pulling the contents out of the container and putting them into the right slots.
This reframing changed the technology approach. Instead of converting PDF to text and then searching the text, we built a pipeline that combined three techniques: layout analysis, named entity recognition, and schema mapping.
Layout analysis identified the structural regions of each document page: headers, tables, form fields, paragraphs, and footnotes. Commercial property documents follow loose conventions. Loss runs are tabular. ACORD applications have labelled fields. Statements of values are a mix of tables and free text. The layout analyser did not need to read the text. It needed to identify which regions of the page were tables, which were form fields, and which were paragraphs.
Named entity recognition operated within each identified region. In a table region, the NER system identified column headers and row values. In a form field region, it identified the label-value pairs. In a paragraph region, it identified insurance-specific entities: dollar amounts, dates, policy numbers, carrier names, occupancy types, construction classes. The NER system was trained on a corpus of eight thousand labelled pages that the firm’s senior underwriters had annotated over a twelve-week period.
Schema mapping connected the extracted entities to the underwriting system’s data model. The ACORD application might list the building address as “Location 1: 450 Main Street, Suite 200, Hartford, CT 06103.” The underwriting system expected separate fields: street, suite, city, state, zip. The schema mapper split compound values, normalised formats, and validated ranges. If a square footage value exceeded five million, the schema mapper flagged it for review because no single building in the firm’s portfolio was that large.
This diagram requires JavaScript.
Enable JavaScript in your browser to use this feature.
The Confidence Score Problem
Extraction was never one hundred percent accurate. The NER system might misread “B” as “8” in a policy number. The schema mapper might split an address incorrectly if the broker used an unusual format. The layout analyser might misclassify a table as a paragraph if the PDF had been scanned at low resolution.
The firm needed to know when to trust the extraction and when to verify it. A fully automated pipeline that was ninety-five percent accurate would produce forty errors per week across eight hundred submissions. In insurance underwriting, an error in the occupancy type or the coverage limits could result in a mispriced policy, which could cost the firm hundreds of thousands of dollars in unexpected claims.
The solution was a confidence score for each extracted field. The confidence score combined three signals: the NER system’s own confidence for the entity recognition, the layout analyser’s confidence for the region classification, and a domain validation check against expected ranges and formats. A field with all three signals above their respective thresholds received a high confidence score. A field with any signal below threshold received a low confidence score and was flagged for human review.
The thresholds were set conservatively. The firm preferred to flag a field for review and be wrong (false positive) than to let an error through and be wrong (false negative). The review rate, the percentage of fields flagged for human verification, was set to approximately twenty-five percent. This meant that underwriters still reviewed one in four extracted fields, but they were reviewing targeted fields rather than extracting everything from scratch. The extraction-plus-review process took twelve minutes per submission on average, compared to the sixty-minute manual extraction that it replaced.
The Automation Boundary
The firm drew a clear automation boundary. Fields with high confidence scores were written directly into the underwriting system without human review. Fields with low confidence scores were presented to the underwriter with the original document region highlighted, the extracted value shown, and a reason for the flag. The underwriter confirmed the value, corrected it, or rejected the extraction entirely.
This boundary was adjusted over time. In the first month, the review threshold was set conservatively and the review rate was thirty-two percent. As the team gained confidence in the system and the error rate in auto-accepted fields remained below one percent, the threshold was loosened and the review rate dropped to eighteen percent by month six. The target was fifteen percent. Enough review to catch systematic errors without burdening the underwriter with trivial confirmations.
The boundary also defined what the system would not attempt. Supplemental questionnaires with free-text narrative responses, “describe the fire protection systems in detail”, were not extracted automatically. The NER system could identify that a fire protection description existed but could not reliably parse the unstructured narrative into structured fields. These sections were left for manual review. Trying to automate narrative extraction would have added complexity without proportional value because the narrative sections accounted for only ten percent of the data points that the underwriting system required.
The Outcome
After six months, the average extraction time per submission dropped from sixty minutes to twelve minutes. The fourteen underwriters reclaimed approximately three hundred and twenty hours per week collectively, equivalent to eight full-time positions. The firm did not reduce headcount. Instead, the underwriters redirected the reclaimed time towards risk assessment, broker relationship management, and complex submission analysis. The underwriting team’s submission-to-bind ratio improved from eighteen percent to twenty-four percent because underwriters could evaluate more submissions and spend more time on the ones that required judgment.
Extraction accuracy for auto-accepted fields was 98.3 percent. The error rate for the prior manual process was estimated at four to six percent based on a quarterly audit sample. The automated system was more accurate than the humans it replaced, which was an uncomfortable finding for the underwriting team but an expected one. Machines are better at consistent data entry than humans who are bored by data entry.
The most significant impact was on the firm’s capacity. Before automation, the underwriting team could process approximately eight hundred submissions per week. After automation, with the same headcount, the team could process twelve hundred submissions per week because the extraction bottleneck was removed. The firm used the additional capacity to expand into two new states without hiring additional underwriters.
When Document Automation Works
Document automation works when the output structure is known in advance. The underwriting system had a fixed data model. The extraction pipeline mapped documents to that data model. The mapping was the specification. If the output structure is unknown or varies with every document, automation is premature. You need to define the target before you can automate the mapping.
Do not start with extraction accuracy. Start with the downstream system’s data model. What fields does it need? What formats do they require? What validation rules apply? Define the target first, then work backward to the extraction requirements. If a field cannot be defined in the target system, do not extract it. If a field can be defined but is never accurate in the source documents, extract it and flag it for review.
The confidence threshold is a business decision, not a technical one. Set it based on the cost of a false positive (unnecessary human review) versus the cost of a false negative (an error that reaches the underwriting system). In insurance, false negatives are expensive. Set the threshold to minimise false negatives and accept the higher review rate. The review rate will decrease over time as the system improves, but the error tolerance should not increase just to reduce review volume.