I built a workflow that collects regulatory filings, extracts comparable product information, and delivers it in workbooks and a dashboard. It supported a $1M engagement and became tooling a national team uses and builds on.
Analysts needed to compare competitors’ products across states. The source material was public, but collecting it meant navigating a stateful portal and reading documents filing by filing. The opportunity was to turn that repeated research into a process the team could run for a new carrier, state, or product line.
I built a phased pipeline that collects metadata and attachments, extracts product provisions into structured records, and produces an Excel comparison workbook and a filterable dashboard. Cloud runs use AWS EC2, S3, and Systems Manager; a separate one-click local package lets colleagues run the tooling on locked-down Windows work machines.
The intelligence supported the sale and delivery of a $1M engagement. A national team adopted the tooling and began building further work on top of it. The commercial result belongs to the engagement; the software made the underlying research repeatable and available to the team.
Collection is resumable: reruns skip completed work instead of starting over. Extracted records follow a strict schema and retain document and page references. A second-model review normalizes terminology across states, and enrichment instructions require a quoted source for each claim. These checks support analyst review; citations give the reviewer a route back to the original material.
Python, Playwright, pdfplumber, openpyxl, AWS EC2 / S3 / Systems Manager, Claude and Codex agents, pytest.