July 2026 · Legal operations

Evaluating a live AI workflow: the release gate

Nobody edited the prompt and nobody changed the vendor, and yet the clause-review step stopped catching what it had caught for months, because the model behind the same product name had moved. A pilot metric is taken once and filed, so it will never show you that; a release gate, run again on every change, will.

A live AI workflow passing through a release gate after model, prompt and connector changes.

How to measure a legal automation pilot, published here in July 2026, answers one question once: did this work. Baseline, error tolerances per class of mistake, a kill-or-scale threshold fixed in advance. Then it stops at go-live, which is where the awkward part starts: a workflow that is actually running keeps asking something else every few weeks, is this still the thing we approved. In between sit all the model versions, prompt edits, retrieval refreshes and connector upgrades that nobody announced.

The ground moves on its own, and somebody has bothered to measure it. Lingjiao Chen, Matei Zaharia and James Zou compared the March 2023 and June 2023 versions of GPT-4, and on the narrow job of telling prime numbers from composite ones accuracy fell from 84 percent to 51 percent. The service name never changed.

So the thing to gate is the change, not the workflow. Build a golden set out of your own matters, with the right answer recorded by the lawyer who would otherwise be doing that work, and start it small: Andrew Ng is right that five examples growing over time beat a complete set that never ships. Run it on every model, prompt, retrieval or connector change, and give it the power to stop one, so that if the critical error class you defined back at the pilot moves at all, the change doesn’t go out. Keep sampling after go-live as well, because the gate only ever sees what somebody put in the set. One named reviewer, a fixed number of live outputs, every week.

Two positions here that people argue with, and I hold both anyway. When a golden set passes at 100 percent for three months running, I stop reading that as a safe workflow and start suspecting a stale set; feed it the failures the sampling found. The second is about vendors, whose own evaluation does not do a gate’s work. The first preregistered study of commercial legal research tools, by Magesh, Surani, Dahl, Suzgun, Manning and Ho (Journal of Empirical Legal Studies, 2025), found the LexisNexis and Thomson Reuters tools then in market hallucinated between 17 and 33 percent of the time, against provider claims the authors call overstated.

None of this is new engineering. The voluntary NIST AI Risk Management Framework already files change management inside its post-deployment monitoring subcategory, MANAGE 4.1. The concept is not what legal teams are short of. Ask who signs off on a model change and you tend to get a pause, then a name that turns out to belong to nobody in particular. That is the gap: a person whose approval the change has to have, and a version number on the set that let it through.

Sources

Next: AI governance for in-house legal teams: five minimum controls →