July 2026 · AI governance

What generative AI actually changes for in-house legal teams

When our in-house team built generative AI tooling on AWS Bedrock, the first real decision was not which model to use. It was which single workflow to point it at first, and in what order to add the rest. We started by writing down what the system must never do before what it could do, then picked one narrow, high-volume task to run first. That sequencing choice taught me more than any product demonstration. The useful question is not whether generative AI will automate a legal department away or leave it untouched, but a narrower one: what does it change, what does it leave untouched, and how do you run it so it helps rather than quietly creating risk.

What does not change

Start with what stays put, because it is the part people skip. The duty to reach a defensible answer does not move. A model can draft a position, but a lawyer still owns it, signs off on it, and answers for it if a regulator asks. Privilege, confidentiality and conflicts do not relax because a tool is in the loop; if anything they get harder, since a poorly configured system can move sensitive material to places it should never be. How far privilege even reaches depends on the jurisdiction: in-house communications are not protected everywhere, and in EU competition matters, for instance, they may attract no privilege at all, so a tool that copies such material into new systems only widens the exposure.

Judgment about risk appetite does not change either. Deciding how much regulatory ambiguity a business can carry, and on what terms, is a matter of context, relationships and history that no model holds. The same is true of the harder calls: when to escalate, when to say no, when to say yes with conditions. Those decisions are the actual job, and generative AI does not touch them.

It is worth saying this plainly, because the point gets lost in product demonstrations. A regulator does not accept "the tool produced it" as an answer, and neither does a board. The lawyer is still accountable. That fact should shape every design choice that follows: the question is not whether to use these tools but how to place them so they support an accountable human rather than obscure one.

What actually changes

What does change is the work that surrounds those decisions, and there is a great deal of it. Three areas stand out in practice.

The first is intake. A large share of an in-house team's time goes to routing questions, restating them clearly, and gathering the facts needed to answer. A well-scoped model can triage a request, ask for the missing pieces, and hand the lawyer a clean brief instead of a vague email. That does not decide anything. It removes the friction before the deciding starts.

The second is knowledge retrieval. In-house teams sit on years of prior advice, playbooks and precedent that are hard to search when they are needed. Retrieval over that material, done carefully, means a lawyer can find how the team handled a similar question two years ago in seconds rather than hours. The value is not a generated answer. It is fast access to the team's own reasoning.

The third is first drafts. Standard clauses, routine responses, summaries of long filings and initial cuts of internal memos are all work where a model can produce a serviceable starting point. The lawyer then edits, corrects and takes ownership. When we built generative AI tooling on AWS Bedrock, the first-phase target was a meaningful, measured reduction in manual workload on exactly this kind of task, together with less routine dependence on outside counsel. The point was never to remove the lawyer. It was to move the lawyer's time toward the work that needs a lawyer.

How to run it responsibly

The gap between a useful tool and a liability is almost entirely in how it is run. A few rules have held up.

  • Set guardrails before capability. Decide what data the system may touch, where outputs may go, and what it must never do, and enforce that in configuration rather than in a policy document nobody reads.
  • Tie review to risk. High-impact work and anything that leaves the team gets a lawyer's review before it goes out; low-risk internal steps can carry a lighter touch. Treat that review as the operational approval gate, not the whole of responsibility: accountability runs across the design, the configuration and the sign-off, not one moment at the end.
  • Run measurable pilots. Pick one workflow, define what good looks like before you start, and measure it. A pilot that cannot show whether it saved time or introduced errors has told you nothing.
  • Track errors honestly. Models are confidently wrong in ways that look plausible. You need a way to catch that, and a team that treats a wrong answer as a finding rather than an embarrassment.

One more point on outside counsel, because it is where the economics are supposed to show up first. Much of what a legal team sends outside is volume rather than hard law: routine research, first-pass review and drafting that goes out because there is no internal capacity, not because the question needs a specialist. Moving a share of that in-house is what teams aim at when they target lower outside-counsel spend. It is a target, not a guarantee; many teams have not yet seen those savings land, while the genuinely difficult matters still go to the firms that should handle them.

None of this requires believing the hype, and none of it requires dismissing the technology. Generative AI does not change what a legal team is for. It changes how much of the surrounding work has to be done by hand, which is worth a lot if you are honest about where the judgment still has to sit. The teams that get value from it stay skeptical about the claims, specific about the workflow, and patient enough to prove each step before trusting it with the next.

Related: Legal tech and AI workflow automation

Next: AI governance for in-house legal teams: five minimum controls →