Read the case study for the business story and results.

The reference labels were the first system we rebuilt.

We reviewed ~50 cited vendor errors with three researchers in a two-hour session after teaching the labeling approach. Decisions used majority agreement. ~60% of those cases were label errors and ~40% were tool errors; this was a diagnostic review of that set.

The annotation guide included positive definitions, paired accepted/rejected examples, ambiguity rules, and a cannot-determine category. Trained annotators applied the guide while scientists authored and adjudicated the standard. We regenerated working imagery from the instrument originals on the client's hardware.

The pipeline separated three decisions.

The stages covered field-of-view quality, cell segmentation, and candidate selection. We built and tested them sequentially against pre-agreed stage thresholds, then measured the combined workflow.

Version-pinned containers and experiment lineage supported reproducibility. The production system ran on the company's own hardware, replacing the hosted computer-vision subscription.

The original test design had limits we later disclosed.

The original training/test split separated image tiles rather than whole slides, plates, or acquisition batches. Related images weren't fully independent. We also lacked an expert-agreement measurement by image-density group.

The field-quality threshold needed a comparison with a trivial predictor. The segmentation score measured pixels, so it couldn't be treated as an instance-level cell score. Candidate-selection recall also needed its operating point and paired precision to describe the tradeoff.

We disclosed these issues to the client, whose team reported remedying them. The revised evaluation design calls for grouped splits, stratified expert-adjudicated references, simple baseline comparisons, and an end-to-end acceptance measure. We didn't independently re-audit their later changes.

Researchers retained control during the shadow period.

The operating process allowed a hold after the model's selection so researchers could review its decision. That enabled shadow operation alongside the existing scientific workflow.

We recommended permanent rescreening of 10% of negatives using random and enriched samples with seeded known positives, plus quarterly blind adjudicated sets. The company budgeted for those checks. This was a research-use system, not a clinical device.

Ownership included the code, data, and operating knowledge.

Six engineers received documentation and walkthroughs. Eighteen months later, the company had extended the workflow to another molecule class without hiring AI specialists.

The delivery eliminated ~$200K in annual vendor subscription expense. The $58M figure on the case page describes first-year platform launch potential. The five-month internal effort's $590K cost had already been incurred and isn't included as a saving.