July 2026 · Legal operations

How to measure a legal automation pilot

A pilot without metrics is a demo. It looks impressive in the room, everyone nods, and three months later no one can say whether it worked. I have run enough of these in a large in-house team to know the plan matters more than the tool.

Start by baselining the workflow before the model touches it. Time the cycle end to end. Count the touches: how many people handle a task and how often it changes hands. Total external spend: outside counsel hours and per-matter fees. Treat these three as the minimum baseline, with quality metrics, accuracy, rework, and error severity, beside them.

Then define the error types each task produces. I sort them into three: critical errors that change a legal outcome, material errors that need real rework, and formal errors that are cosmetic. Set the tolerance per class, zero for the critical one, and write it down before the pilot.

Measure reviewer time, not model time. The model answers in seconds; what matters is how long a lawyer checks and fixes it. If review erases the saving, the pilot failed however fast the model ran.

Decide the kill-or-scale threshold in advance. A widely discussed 2025 report from MIT NANDA, The GenAI Divide, found that 95 percent of enterprise generative AI pilots showed no measurable impact on profit and loss. In its 2025 report, Thomson Reuters found only 20 percent of respondents across professional services say their organization measures the return on AI. If you are not measuring, you cannot know which side of it you are on.

Baseline first, set error tolerances, measure reviewer time. Then capture it in one place: a pilot decision record naming the owner, date, kill-or-scale threshold, and the evidence behind the call. That record, not the applauded demo, tells you whether to scale.

Sources

Related: GenAI transformation for legal teams

Next: The billable hour is running out of logic →