The question nobody wrote down before go live
A contract review system has been live for four months at a mid-market firm. It flags indemnification caps, non-standard termination language and jurisdiction clauses across a few hundred contracts a month, and the associates who used to do that first pass now review its output instead. It works. Then the clause library needs a new entry for a client's updated data processing addendum, or the model provider ships a new version, or someone on the team wants to tighten the prompt because it keeps missing a specific carve out in indemnification clauses. Who signs off on that change? Where is it written down that they did? And when a client's outside auditor asks six months later why the system flagged a clause differently in March than it does now, who answers, and from what record?
This is the part of AI contract review change control that almost never gets built before go live, because it is not a data privacy question or a model risk question in the abstract. It is a mechanical one: three different kinds of change, each needs a different owner, and none of it gets decided while everyone is focused on shipping the thing in the first place.
Three changes, not one
Teams that have not run a production legal AI system before tend to treat "changing the system" as one event. It is not. A contract review system in production has three distinct change surfaces, and treating them the same way is how firms end up either approving everything at partner level, which is too slow, or approving nothing formally, which is how a system drifts without anyone noticing.
The clause library. This is the firm's own reference set: the standard positions, the acceptable ranges, the red flag definitions that tell the system what counts as a non-standard termination clause or an uncapped indemnity. It changes often, usually because a partner wants a new client position reflected or a precedent has shifted. This is legal judgment. It belongs with whoever already owns clause positions in the firm's playbook, typically a senior associate or partner in the relevant practice group, not with whoever runs the AI system.
The prompt and the review logic. This is the instruction set that tells the system how to read a contract, what to extract, how to phrase a flag, and when to escalate rather than answer. A small wording change here can shift which clauses get flagged and which get missed, in ways that are not obvious from reading the prompt itself. This is an engineering and legal judgment problem together, and it needs both in the room, because a change that reads as a harmless clarification can quietly change the system's recall on a clause type nobody tested.
The model version. The underlying model gets updated, sometimes by the firm's own choice, sometimes because the vendor deprecates the old version and forces the move. This looks like the smallest change, since the prompt and the clause library stay the same. It is often the one that moves behavior the most, because a new model version can interpret the same instructions differently.
Each of these has a different blast radius, a different kind of error, and should have a different approver. Routing all three through the same sign off either slows the firm down on the changes that need to move fast, like a clause update for an urgent client matter, or waves through the changes that deserve the most scrutiny, like a model swap.
What approval actually requires, not just who
A sign off that is just a name on an email does not survive an audit. What a client's compliance team or an outside auditor wants to see is a record that shows four things for any change: what changed, who approved it, what evidence they reviewed before approving, and what happened to the system's output before and after.
That means every change of consequence needs a before and after comparison on a fixed set of test contracts before it goes live, not after. This is the same discipline a firm would use for any evaluation set: a standing collection of real, anonymized contracts with known correct answers, run through the system before a change is approved so the approver sees actual output, not a description of the change. For a prompt change, that might mean twenty contracts covering the clause types the firm cares about most. For a model version update, it means the same set, because a vendor's release notes describing general improvements say nothing about whether this firm's specific clause types are affected.
Without that test set, approval becomes a judgment call based on how the change was described rather than how the system now behaves, and that is exactly the gap an auditor will find.
What happens when a change goes wrong
The honest answer to "who approves" has to include "what happens when the approver got it wrong." A clause library update that was too aggressive might start flagging standard language as non-standard, generating noise that associates learn to ignore, which is its own kind of failure. A model update might quietly reduce recall on a clause type the firm rarely sees, so the gap does not surface until a contract with that clause slips through unreviewed.Both of these point to the same requirement: the system needs to keep a log of what version of the prompt, clause library and model produced any given output, tied to the date. Without that, when an error surfaces weeks later, there is no way to tell which version produced it, whether it was already fixed, or whether the same error is still live. This is the detail that turns change control from a sign off ritual into something that actually protects the firm. A related problem, the gap between a demo working well and a system holding up once it is running on contracts nobody hand picked, is covered in why your AI pilot demoed well and never shipped.
What has to exist before any of this works
None of the above is possible without two things already in place, and firms that skip them end up improvising change control after something has already gone wrong.
The first is the test set itself: a standing, maintained collection of real contracts with known correct flags, kept current as the firm's clause positions evolve. This is not a one time build. It needs an owner who adds new contract types as the firm's work changes, or the test set goes stale and stops catching the failures that matter.
The second is a version record built into the system from the start, not bolted on later. Every output needs to be traceable to the exact prompt version, clause library version and model version that produced it. Retrofitting this after a system has been live for a year, once outputs from a dozen different silent configurations are already sitting in client files, is considerably harder than building it in from day one.
The decision this puts on your desk
If you own this system, the decision is not whether to have change control. It is whether you assign three separate approval paths, one for clause library changes, one for prompt and logic changes, one for model version changes, each with its own owner and its own evidence requirement, or whether you leave it as one informal process that will not hold up the first time a client's auditor asks for the record. The second path is the one most firms are currently on, usually without having chosen it on purpose.
Where we come in
Alppoint builds contract review systems with the version record and the evaluation set designed in from the start, so change control is a workflow the system supports rather than a policy bolted on after launch. We do this as the team building the system end to end, or alongside a firm's existing legal technology and AI staff on the specific workstream of getting change control production ready. Either way, the system, the test set and the approval record stay the firm's own, running in its own environment.



