AI in Finance

AI for KYC screening: where it earns its keep in finance compliance

Published 4 May 2026

KYC screening is one of the use cases I expect to land in production fastest with the new generation of finance agents. The volume is high, the inputs are largely textual, the structure of the work is well understood, and the cost pressure on compliance functions has been severe for years.

The KYC screener was one of the ten templates Anthropic shipped in the May launch. It is also one of the templates I would treat with the most rigorous governance, because compliance does not allow for “the agent got it nearly right.”

This post is the practitioner read on where a KYC screening agent earns its keep, where the risk sits, and what a compliance function should require before putting one in front of real work.


What the screener actually does

In the Anthropic template, the KYC screener assembles entity files, reviews supporting documents, and packages escalations for human reviewers.

What that looks like operationally is a multi-step pipeline. The agent ingests the new client or counterparty data. It pulls supporting documents from connected sources. Dun & Bradstreet is the named data partner Anthropic published alongside the KYC template, providing the D&B Commercial Graph and D-U-N-S Number for entity identity. (Source: anthropic.com/news/finance-agents.) Moody’s Compliance Catalyst, Verisk, and the existing PEP and sanctions feeds are available through the wider Claude for Financial Services connector ecosystem rather than as the screener’s named default. The agent cross-references the inputs against the lists and risk indicators the compliance function cares about and produces a structured file with a recommended risk rating and the supporting evidence attached.

The human in the loop is the reviewer who accepts, escalates, or rejects the file.

That is a real productivity gain for a function that has historically spent most of its time on the assembly and the cross-reference, not on the judgment. It is also a real risk if the assembly is treated as the judgment.


What the screener gets right

Three things.

The volume. A trained reviewer doing KYC manually can clear a small number of files per day. An agent doing the assembly clears a much larger number, and the reviewer who used to do both is now free to focus on the escalations. The volume gain is the headline benefit.

The consistency. Manual KYC has variable depth depending on who is doing it, what shift they are on, and how busy the queue is. An agent applies the same procedure to every file. That consistency is itself a compliance asset, because it produces an audit trail the regulator can interrogate cleanly.

The packaging of the escalation. The escalation that lands in front of the human reviewer is structured, evidenced, and has the supporting data attached. That is a meaningful upgrade on the typical manual escalation, which is often a flag with a free-text comment.


What the screener gets wrong, or will

Compliance has zero tolerance for false negatives in some categories and zero tolerance for unjustified false positives in others. The agent will produce both. The industry baseline for transaction-monitoring and adverse-media screening false positives is 90 to 95%, with best-in-class programmes running at 30 to 50%. (Source: Flagright benchmark.) That is the floor the agent is being introduced into. The agent will not fix the false-positive rate on its own.

False negatives. The customer who should have been escalated and was not. The PEP whose name is spelled differently in the source than in the watch-list. The corporate structure whose ultimate beneficial owner is hidden behind a holding that the agent has not been trained to unpick. The connector data that is incomplete because the data partner has not updated it. The model that produces a risk score on the data it has, not on the data it should have had.

False positives. The escalation that should not have been an escalation. The legitimate transaction flagged because the agent has been over-conservative on a category that does not apply. The reviewer queue clogged with files the agent did not need to escalate. The cost of those is not a fine. It is the cost of the reviewer’s attention being diluted on noise, which raises the risk of missing the signal.

The compliance function that deploys the screener without designing for both of those failure modes will be the compliance function that finds out about them in a regulator’s letter. The cautionary tale most worth keeping in mind is not about AI specifically. TD Bank paid over $3 billion in October 2024 across DOJ, OCC, FinCEN, and the Federal Reserve, the largest BSA fine ever, for failing to monitor 92% of its transaction volume over six years. (Source: DOJ case page.) There is no published enforcement action yet that names AI-assisted KYC as the cause of the failure. That absence does not mean the risk is theoretical. It means we are early.


What I would require before deploying

The list is short. Most of it is non-negotiable.

Documented decision logic. The skill the agent is following has to be written, versioned, and reviewable by the compliance function. The regulator will ask. The vendor that does not let you see the prompt is the vendor whose agent you cannot defend in audit.

Full audit trail. Every screening decision the agent makes has to be reconstructable. The inputs it saw, the connectors it called, the data the connectors returned, the subagents that were used, the rules that fired, the output it produced. The AI governance framework is the long version. The shorter version is the regulator will want this and you do not get to add it later.

Human review on every output, not every escalation. This is the part finance functions get wrong. The reviewer should be sampling the agent’s clear-pass decisions as well as reviewing the escalations. The clear-pass population is where the false negatives hide. Sampling rate can be lower than the escalation rate, but it cannot be zero.

Defined connector quality contract. What the connector returns, at what cadence, with what coverage. The data partners listed in the launch are credible. The contract still needs to be specific to the screener’s use case. A connector that does not refresh sanctions data daily is not a sanctions connector you can deploy on real KYC.

Escalation paths for agent failure. What happens when the agent’s confidence is low. What happens when the connector is down. What happens when the agent’s output is internally inconsistent. The function needs a documented escalation for each, in advance, and the reviewer needs to know what to do.

Training for the reviewer. The reviewer’s job is now to interrogate the agent’s output, not produce their own. That is a different competence. The function that does not train for it will have reviewers who rubber-stamp the agent’s recommendation, which is the worst of both worlds.


What changes about the compliance function

The shape of the compliance team changes when the screener lands in production.

The volume layer shrinks. The reviewers who used to do file assembly are doing escalation review and sampling instead. Some of them are doing the more substantive part of the work for the first time.

The senior layer grows in importance. The compliance head who used to manage the workflow now spends more time on the governance of the agent and the sampling regime. The conversation with the regulator is about the design of the AI-assisted process, not the productivity of the human one. That is a different conversation, and the compliance leaders who handle it well will be visibly more valuable inside their organisations.

The recruitment profile changes too. The compliance function still hires juniors. The juniors hired now will start in the agent-assisted environment. The competences they develop will be different from the competences the current senior team developed. That is not a problem to solve. It is a transition to design for, the same way every other part of the function is.


What I would do with this

If I were running a compliance function this year, I would pilot the KYC screener in Q3 with the head of compliance co-owning the rollout. I would run it in parallel with the manual process for three months. I would not switch the workflow over until the parallel data showed the agent and the manual process agreeing on the high-risk decisions, with the disagreements understood category by category.

I would invest in the reviewer training in advance of the deployment, not in response to it. The agent is only as good as the people interrogating its output.

I would have the audit trail design reviewed by external counsel before the deployment, and I would document the model and the deployment for the regulator on day one rather than waiting for them to ask.

The AI vendor evaluation framework applies here. KYC is one of the use cases where the questions in that framework matter most.


Where this lands

KYC screening is the use case where finance AI agents look most ready, and also the use case where the cost of a wrong deployment is highest. The two are not in tension. They are the same point. The use cases that promise the most are the use cases that demand the most discipline in the deployment.

The compliance function that lands the screener well will spend less time on file assembly and more time on the substance of the risk. The compliance function that lands it badly will be the case study the regulators cite.

Choose which one yours is going to be.


Maebh Collins is a Fellow Chartered Accountant (FCA, ICAEW) with Big 4 training and twenty years of operational experience as a founder and senior finance leader.

Back to Blog | AI in Finance →