July 2026 · Legal operations

Evaluating a live AI workflow: the release gate

Nobody edited the prompt and nobody changed the vendor. The clause-review step simply stopped catching what it had caught for months, because the model behind the same product name had moved. A pilot metric cannot see that. A release gate can.

How to measure a legal automation pilot, published here earlier this month, answers one question once: did this work. Baseline, error tolerances per class of mistake, a kill-or-scale threshold fixed in advance. It stops at go-live. A live workflow asks a different question every few weeks: is this still the thing we approved. In between sit all the model versions, prompt edits, retrieval refreshes and connector upgrades nobody announced.

The ground moves on its own. Lingjiao Chen, Matei Zaharia and James Zou compared the March 2023 and June 2023 versions of GPT-4: on telling prime numbers from composite ones, accuracy fell from 84 percent to 51 percent. The service name never changed.

So gate the change, not the workflow. Build a golden set from your own matters, with the right answer recorded by the lawyer who would otherwise do that work; Andrew Ng is right that five examples growing over time beat a complete set that never ships. Run it on every model, prompt, retrieval or connector change and let it block: if the critical error class you defined for the pilot moves at all, the change does not ship. Sample after go-live too, because the gate sees only what is in the set. One named reviewer, a fixed number of live outputs, every week.

Two positions worth disagreeing with. A set that passes at 100 percent for three months is not a safe workflow, it is a stale set; feed it the failures the sampling found. And a vendor's own evaluation is not a gate: the first preregistered study of commercial legal research tools, by Magesh, Surani, Dahl, Suzgun, Manning and Ho (Journal of Empirical Legal Studies, 2025), found the LexisNexis and Thomson Reuters tools then in market hallucinated between 17 and 33 percent of the time, against provider claims the authors call overstated.

None of this is novel engineering. The voluntary NIST AI Risk Management Framework already files change management inside its post-deployment monitoring subcategory, MANAGE 4.1. What legal teams lack is not the concept but the owner: someone whose sign-off the change needs, and a version number on the set that let it pass.

Sources

Next: AI governance for in-house legal teams: five minimum controls →