Access checks happened before documents reached the model.
We used their existing matter entitlements to determine which documents an attorney could retrieve. Their security team connected that authorization system to the new AI system.
This preserved one authority for access. Maintaining a second set of permissions would have created another place where changes could drift from their own records.
The retrieval process applied those permissions before selecting documents for the model. The resulting answer could draw on the material the attorney was entitled to access under their policies.
The test had to answer with what the deal team knew at the time.
We built test sets of ~150 to 200 questions per workflow from their historical matters.
For each test, we withheld the test matter's own documents and limited retrieval to information that predated the deal. This prevented the system from appearing capable simply because the finished answer was already in the material it could search.
The evaluation combined automated checks, a model-based judge, and sampled human review. We reran the question sets when models or prompts changed. We also tested whether the system could expose material across matter-access boundaries.
A citation had to lead somewhere real.
Before an AI answer passed the citation checks, the source had to exist, the reference had to parse, the cited location had to resolve, and quoted text had to match the source.
Those checks handle concrete failures that software can identify reliably. A partner still reviewed whether the cited material supported the legal conclusion. When confidence was too low, the system could abstain instead of producing an unsupported answer.
Together, source checks, abstention, and partner review made the AI answer easier to assess while preserving the attorney's role in the work.
We changed where we put the engineering effort.
We initially gave practice-specific fine-tuning too much weight. Evaluation showed that verified retrieval and the citation checks were doing the work that improved grounding.
We made those controls central to the architecture. Fine-tuning became a selective tool for stable house formats, using abstracted and redacted material. That also reduced the need to retrain separate models with every base-model change.
Matter data needed a full lifecycle.
We treated the retrieval store as privileged data. It used per-matter encryption keys, deletion procedures covering stored vectors, and a backup strategy tied to those keys.
A litigation-hold check sat in the deletion process. Closing a matter didn't automatically mean its data could be destroyed. The workflow checked hold status before carrying out destruction.
We measured useful output and usage separately.
| Measure | Definition and collection |
|---|---|
| Rework: 83% → 24% | Share of AI outputs needing edits or a redo. The same one-week instrumented study ran in their corporate practice before the build and ~3 months after launch, with the same population and method. |
| Usage: 32% → 72% | Share of attorneys using AI in their corporate practice. The earlier value came from the previous vendor's analytics; the later value came from our telemetry. Use was mandated before and voluntary after. |
| Cloud-AI expense eliminated: ~$1.7M/year | Their recurring cloud-AI invoice bundle, ~$141,000 per month, that was removed. |
The observation and processing for the time study stayed on their servers. Reports removed case details and grouped activity by work type. Our team received those reports rather than raw captures of privileged matters.
Handover included operation, changes, and diagnosis.
We trained their knowledge-management team and named champions to create citation-bound workflows through the no-code layer. We also taught them to review past calendars and study live work weeks to identify further bottlenecks.
The operating capability covered the infrastructure, audit records, alerts, and model-change process. Changes could be staged, checked against the test questions, promoted, and rolled back.
Their team used that capability to extend the system beyond M&A into other practice areas. Specialist work remains part of maintaining quality: interpreting evaluation failures, adjusting when the system should abstain, and assessing new models or unfamiliar threats.