Build | October 3, 2026

Where the human sits in human in the loop AI for transaction monitoring alerts

Generic content on human in the loop AI stops at definitions and citations. This post goes inside an AML alert review queue to show what the model surfaces, what the investigator checks, and what gets logged when they overrule it.

Where the human sits in human in the loop AI for transaction monitoring alerts

The alert that lands on a Tuesday morning

An investigator opens the queue and there is a new alert: a wire transfer that trips a velocity rule, three transactions in two days against an account that usually moves money once a month. The old process had the investigator pull the account history, check the customer's stated business, look at counterparties, and write up a disposition. Twenty to forty minutes, most of it spent confirming the alert is nothing. This is the workflow every transaction monitoring team runs, and it is also the workflow that gets named whenever someone proposes AI for AML and the compliance leader in the room asks a fair question: who actually looks at this before it counts?

That question is where most writing on human in the loop AI stops short. It will tell you the concept exists on a spectrum, cite a regulation, and move on without saying what the human in the loop is actually looking at, what they can see that the model cannot, and what happens to the record when they disagree with it. For a transaction monitoring program, that is the only part that matters, because it is the part an examiner will ask about and the part that determines whether the system reduces investigator time or just moves the risk somewhere less visible.

What the model actually does to the alert

Before any of this works, the alert has to arrive with something attached to it beyond the triggering rule. A raw alert from the monitoring engine tells the investigator a threshold was crossed. What the AI layer adds is a case file: the customer's transaction history over a relevant window, the stated business purpose from onboarding, prior alerts and how they were disposed, counterparty patterns, and a narrative that ties them together in plain language. "Three transfers totaling $340,000 over two days, to a counterparty not seen on this account before, in a customer whose stated business is retail goods import. Two prior alerts on this account were closed as false positives for seasonal volume."

That narrative is a draft, not a finding. It is built from retrieval against the bank's own transaction and case history, not from the model's general knowledge, and it is built so every claim in it traces back to a specific record the investigator can open. The system also assigns a preliminary risk read, something like "pattern consistent with prior false positive" or "counterparty and amount both new for this account," because that framing is what tells the investigator where to spend their attention first.

What the investigator checks before disposing the alert

This is the part a workflow post has to show, because it is the part that proves the human is doing something other than rubber stamping a model output. The investigator does not re-run the whole investigation from scratch, but they do not accept the narrative on its face either. In practice they check three things every time.

  • The source behind each claim. If the narrative says the counterparty is new to this account, the investigator opens the transaction record it was pulled from and confirms it. This is a few seconds per claim, not a re-investigation, but it is the step that catches a model that summarized the wrong account or missed a transaction outside the window it was given.
  • What the narrative left out. A model trained to summarize patterns can miss context a person would weigh, like a customer service note about a recent business change, or a prior SAR filed on a related account that sits in a different system the alert pull did not reach. The investigator's judgment call is whether the case file is complete enough to decide on, not just whether it is accurate as far as it goes.
  • Whether the disposition the system suggests matches their own read. The system may suggest "close, consistent with known business pattern." The investigator either agrees, or they do not, and either way they are the one who closes the alert, files the SAR, or escalates it.

What this buys the team is not a faster rubber stamp. It is a faster version of the same judgment the investigator was already applying, because the fifteen minutes that used to go into assembling the case file now goes into checking and deciding. The investigator spends less time being a researcher and the same amount of time, or close to it, being a decision maker.

What gets logged when the investigator overrules the model

The record that gets kept is the part regulators and internal audit actually test, and it is the part generic content on this topic skips entirely. Every alert disposition carries three things on the record: what the system surfaced and suggested, what the investigator decided, and, when the two disagree, why. If the model suggests closing an alert as a known false positive pattern and the investigator escalates it anyway because the counterparty also showed up on an unrelated sanctions screening hit, that reasoning gets written down, not just the outcome.

This matters for two separate reasons. The first is day to day quality control: a pattern of overrides in one direction tells the program whether the model is systematically too conservative or too lenient on a particular typology, and that is the input that drives retuning, not a guess. The second is the audit trail a program needs to produce when an examiner or an internal reviewer asks why a given alert was closed the way it was, six months after the fact. "The system suggested X, the investigator disagreed and did Y because Z" is a defensible record. "The system suggested X and it was closed" is not, regardless of whether the decision was actually right.

This is also the point where ai guardrails earn their keep in a way that is concrete rather than conceptual. A guardrail here is not an abstract safety layer, it is a specific rule: the model cannot auto close any alert above a dollar threshold, cannot dispose of an alert tied to a prior SAR without investigator sign off, and cannot change the risk rating it already assigned to a customer without a second review. Those are decisions a compliance leader makes about where the firm's own risk appetite sits, and they are what turns "human in the loop" from a design philosophy into a set of enforceable rules the system actually runs under.

Where this fits NIST AI RMF and why that matters for alert review specifically

The nist ai rmf framework is cited constantly in AML AI content, usually as a citation and nothing else. What it actually asks a program to do, in terms this workflow can use, is map, measure, manage and govern the system's risk across its lifecycle, which in alert review terms means: know what the model is for and where it is not to be trusted (map), test its outputs against known cases before and after deployment (measure), decide what a wrong suggestion costs and build controls sized to that cost (manage), and keep a standing record of who is accountable for the system's behavior over time (govern). None of that is satisfied by a vendor's description of their model. It is satisfied by the specific logging, thresholds and override record described above, applied consistently across every alert the queue produces.

The same discipline is what makes an ai audit of the alert review process a known quantity rather than an open question. An auditor sampling closed alerts should be able to pull any disposition and see the case file the investigator saw, the suggestion the system made, and the reasoning behind any disagreement. If that trail does not exist for every alert, the program has deployed a model, not a governed system, and the difference shows up exactly when an examiner asks for it. This is also the core of what we mean by ai risk management in this context: not a policy document sitting beside the monitoring system, but a set of rules the system enforces on every alert whether anyone is watching that day or not.

What has to be in place before any of this works

The case file the model builds is only as good as the data it pulls from. If transaction history, prior disposition records and onboarding documentation sit in systems that do not talk to each other, the system is guessing at completeness, which is the exact failure mode the investigator's "what did this leave out" check exists to catch, but a program should not be relying on that check to cover for missing integration. The second requirement is a clean record of past dispositions to build from, since the model's read on "known false positive pattern" is only trustworthy if the history it is drawing on was itself disposed correctly. A program replacing bad judgment with faster bad judgment has not reduced its risk.

Where we come in

Alppoint builds the alert review layer described here directly into a client's own environment and monitoring stack, including the case file assembly, the guardrails around auto disposition, and the override logging an audit actually needs, whether that means owning the build end to end or working alongside a compliance or data team that already has a monitoring system and needs this layer built properly around it. For the governance structure that sits above a single workflow like this one, including ownership, review and risk tracking across every AI system in the firm, see our AI governance work. If your team is weighing where alert review sits against other candidates for AI investment, our post on choosing the first workflow to automate covers how to make that call.

The decision in front of you is not whether to add AI to alert review. It is where the investigator's judgment sits in that review, what gets logged when they exercise it, and what you can show an examiner the day they ask. Get a conversation started with us when you are ready to settle that.

Related posts