Read the case study for the business story and results.
The data foundation came before the interface.
We processed an archive of ~15 years of image-heavy marketing emails, extracted campaign values, and reconciled duplicate provider feeds into consistent records. On the canary evaluation sample, extraction agreed with the company's experts 98.4% of the time.
The rebuilt store supported exact numeric queries. Production retrieval averaged ~50 milliseconds, with three-to-six-second outliers. This measures the retrieval step, not the entire answer workflow.
Exact figures and written explanations followed different paths.
Numeric claims came from deterministic queries against the canonical records. Definitions could use glossary-backed retrieval. The AI composed the explanation from the retrieved evidence and attached sources.
A separate judge routed risky question types to human review. Tests covered ambiguous brands, dates, and questions that the available data couldn't support. The system surpassed our internal 95% groundedness acceptance bar with human escalation.
We tested against the company's own client work.
The evaluation used six months of historical client questions and the answers their experts had provided. Partners compared outputs blind. More than 400 automated tests also covered regression and known-good signal calculations.
During the first two months, more than 97.6% of questions needed no human escalation. This is the share handled automatically, not an accuracy percentage. That window covered the original 27 clients and ~50% of the new clients as they came on board.
We added the historical versioning the first build missed.
The first store represented the current best view. After the engagement, we disclosed that historical values could change when records were reprocessed and delivered the point-in-time correction at no charge.
The revised approach retained historical signal versions so a query could recover the value known at a given date. The lesson was to test historical stability alongside current accuracy.
Their team could operate and extend the foundation.
We taught their team to add providers, adjust escalation boundaries, create new escalation categories, and build citation-backed workflows. Their later client-facing interface was built on the delivered foundation without our implementation work.
The commercial result was ~$3.6M in signed annual recurring contracts with billing started. Reclaimed partner hours were net of the new oversight tasks. Those are distinct measures: recurring contract value, delivery capacity, and earnings.