Run | October 9, 2026

What happens on day two hundred of an AI due diligence review system

Launch day proves an AI due diligence system works. Here is what quietly breaks months later, and the monitoring a firm needs to catch it before a client does.

What happens on day two hundred of an AI due diligence review system

Day one of an AI due diligence review system is easy to picture. The deal team points it at a data room, it classifies the contracts, flags the change of control clauses and the termination rights, and a junior associate who used to spend three days on first pass review spends three hours checking the system's work instead. Everyone is pleased. The partner who sponsored it tells the next deal team to use it too.

Day two hundred looks different, and almost nobody writes about it. The deal team has turned over twice. The data room on the current deal has a seller who exports everything as scanned image PDFs instead of native text, because that is what their document management system does by default. The review categories drift, because the associate running this deal interprets "material adverse change language" slightly differently than the one who ran the deal in March. Nobody has looked at the system's accuracy numbers since the week it launched, because nobody was ever assigned to look. This is the real subject of AI due diligence review monitoring: not whether the system worked at go live, but what it is doing right now, on the deal in front of you, and whether anyone would notice if it quietly started doing it wrong.

What the system does, day to day

Strip away the launch narrative and the mechanics are plain. Documents land in a data room or a shared drive. The system reads each one, classifies it against a taxonomy (employment agreement, IP assignment, change of control clause, indemnification cap, and so on), extracts the terms that matter for the category, and flags anything that trips a defined risk rule, such as a termination right triggered by the transaction itself. It produces a working file: a structured summary per document, a risk flag where one applies, and a link back to the page and clause the flag came from.

That output goes to an associate before it goes anywhere near a partner or a client. The associate's job is not to re-read every document. It is to check the flagged items against the source, confirm the system read the clause correctly, and decide whether the flag is a real issue or a false positive. Unflagged documents get a lighter pass, often a spot check on a sample, not a full re-review, because the whole point of the system is to stop paying senior time for first-pass reading. That division of labor, what the system decides on its own versus what a person has to confirm before it counts, is the thing that has to survive contact with month six, not just week one.

Where AI due diligence review monitoring actually breaks down

Three failure modes show up repeatedly in a system like this, and none of them look like a crisis when they start.

New document formats. The system was tuned on the document types from the deals it launched with: clean native PDFs, a fairly standard set of agreement templates. A new deal brings a seller whose documents are scanned, or translated, or formatted by a data room platform the system has never seen. The extraction quality on those documents quietly drops. Nothing throws an error. The system still produces a summary and a confidence-looking output, it is just wrong more often, and wrong in a way that looks the same as a correct answer until someone checks the source.

Deal team turnover. The person who understood why a particular clause category was defined the way it was has moved to a different deal, or left the firm. The new associate running review interprets the categories a little differently, flags things the system was never tuned to catch, or stops flagging things the prior team always escalated. The system has not changed. The judgment layered on top of it has, and nobody wrote down what the judgment was supposed to be.

Threshold drift. Every review rule in a system like this has a threshold behind it, explicit or not. How material does a change of control clause have to be before it is flagged. How close a termination date has to be to the closing date to count as a risk. Those thresholds were set against the deals the system launched on. As deal size, industry and jurisdiction shift, the thresholds that made sense in month one start producing too many flags, or too few, and the associate reviewing the output starts trusting or distrusting the system based on last week's experience rather than this week's calibration.

None of these show up as an outage. The system keeps running, keeps producing output that looks the same shape as it always has. That is exactly why they are dangerous: a system that fails loudly gets fixed, a system that fails quietly gets trusted.

What has to be in place before day two hundred arrives

Catching drift requires deciding, before launch, what you would check if you suspected the system had gotten worse and nobody had told you. In practice that means three things.

A standing accuracy sample. A small set of documents, reviewed in full by a senior associate on a fixed schedule, independent of whatever deal is live that month. The point is not to re-review everything. It is to have a number that would move if the system's accuracy moved, rather than relying on a partner's gut sense that something feels off.

A named owner for the review taxonomy. Someone who is accountable for what "material adverse change language" means this quarter, who approves a change to that definition the way a change to the firm's own precedent library would be approved, and who is the person a rotating deal team asks instead of guessing. Without this, the categories drift with whoever is holding the pen.

A log of what the system flagged, what the associate did with the flag, and why. This is the record that lets you see threshold drift before it costs you a deal. If associates are overriding flags at a rate that climbs month over month, that is either the system getting worse or the deals getting harder, and the log is the only way to tell which. It is also the record a client or an insurer will ask for if a missed issue surfaces after closing.

This is the same discipline that governs change control on a contract review system once it is live: someone has to own what changes, and there has to be a record of why. We wrote about who holds that approval authority in who approves a change to an AI contract review system after it goes live, and the logic carries straight across to due diligence review.

What happens when the system is wrong

A missed flag in due diligence is not a UI bug, it is a liability exposed to a client months after the deal closed. So the question worth answering before go live, not after an incident, is what the recovery path looks like. If a document was misclassified or a clause misread, who finds out, how far back does the firm look to see if the same error pattern touched other documents on the same deal or a prior one, and what gets disclosed to the client. A system with no standing sample and no change log cannot answer any of that quickly. It can only start looking once someone has already noticed a problem, which in due diligence usually means after the deal has closed.

The decision a firm running one of these systems actually owns is not whether to monitor it. It is who is accountable for noticing when it drifts, on what schedule, and with what authority to pause the system or change the taxonomy without waiting for the next scheduled review. That decision sits with the partner who sponsored the tool, not with whichever associate happens to be running the current deal, and it has to be made explicit rather than assumed.

Where we come in

We build the monitoring layer alongside the review system itself: the accuracy sampling, the taxonomy ownership, the flag and override log, structured so a partner can answer "how is it doing right now" without commissioning a special review to find out. That work can run end to end, as a defined workstream, or alongside a firm's existing data and engineering team. If you have a due diligence review system running past its first few months, we are glad to look at what it is actually producing today.

Related posts

Put frontier AI to work in your firm

A two-week diagnostic tells you whether the problem you have in mind is worth building.