WorkExperimentsServicesReferences
Vibe CodingAgents

TAX DATA EXTRACTOR

Puritas Springs hired me to build a proof-of-concept for an AI-powered data extraction service in the summer of my forty-eighth year. Over the next three days, one of the partners and I built two complementary tools, and then we connected them together.

Puritas sells software used by small law firms to complete complex tax forms, and data accuracy and human review were their highest priorities.

I knew we would need sample documents that represented the variety of forms, photo copies, cellphone images, and digital files that would be typical when attempting even basic form completion, let alone anything involving real estate or international finance. I immediately thought of one family that had both the financial variety and desperately needed levity to get me through AI evals of tax documentation.

Key Results

  • Launched tax-data extraction web service that reliably converts a mixed set of client documents into a structured, reviewable JSON record.

  • Integrated JSON results into existing software with extracted values carrying evidence, missing-data flags, and conflicts so lawyers can review the records before commiting data to forms.

  • Built functional evaluation framework into the service dashboard to allow for observable improvements to prompts, model selection, and cost controls.

I had ChatGPT create a library of documents from Raleigh St. Clair's 1040 PDF to a photo copy of a charitable donation made to Neville Smythe-Dorleac's foundation. Chas Tenenbaum had dispositions of capital assets and Eli Cash had cellphone pictures of Consolidated Tax Reporting Statement. All this was created with deliberate typos, conflicting data, and missing information required by the target forms.

The batches of documents came with Golden Records, example JSON files that detailed what perfect extraction looked like: all available data captured, conflicts resolved, missing data flagged, and links to source documents to aid human review.

The prototype was designed to make model choice a configuration decision, and allow us to test success rates across providers and models. In candidate mode, the system ranks evidence, resolves aliases and conflicts, then makes a second verification call before returning final data.

EXTRACTION SERVICE DIAGRAM

The service worked reliably. It would sometimes miss tricky to extract data, but that was to be expected. We designed the eval test cases to be difficult. The goal was to deliver time savings and minimize hallucinations inside a workflow that kept the human in full control. That was achieved.

Surprising finding: Claude Haiku 4.5 turned out to be the best bang for buck delivering consistently excellent extraction scores at less than 10¢/run. Our conservative estimates put the time saved at 30 minutes of document review and data entry.

Go, Mordecai.

Project Team

Charlie DeMarcoProject Lead
Larry HouselPuritas Partner

More Experiments

BUILDING A PORTFOLIO with AI

BUILDING A PORTFOLIO with AI

© 2026 CHARLIE DeMARCO
All marks and design are copyright of their respective owners.
v5.0.2