AI Agents for Regulatory Compliance: What They Actually Do

In short

The FCA said in September 2025 that it does not plan to introduce extra regulation for AI and will rely on existing frameworks such as the Consumer Duty and the Senior Managers and Certification Regime, so a compliance agent can compress reading, triage and drafting but cannot carry the accountability. A useful agent runs on a schedule, writes tracked records with written reasoning, and leaves a run record that a human can audit a year later.

AI agents for regulatory compliance are software processes that retrieve regulatory publications on a schedule, score them against your specific business with written reasoning, and hand what matters to people as tracked obligations, leaving a run record a human can audit. The demand is not in doubt: the Bank of England and FCA's 2024 survey found that 75% of UK financial services firms already use AI. The open question is what a compliance officer should insist on seeing before trusting one.

A vendor that cannot show you one complete run record is selling a feed with adjectives. Below, a capability and limitation matrix shows, component by component, what an agent reads, what it writes and where a human decides in the workflow of RegWatch, my company. Then come the fields a complete run record should hold and what an auditor should be able to check in it a year later, how failure should be recorded instead of hidden, and six questions to hold any vendor, including us, to the same standard. The AI compliance hub is built around this article.

"AI agents for compliance" is three different questions, so make sure you are asking the right one

Searches for these terms mix three intents, and buying decisions go wrong when teams do too.

1. Agents doing compliance work. Software that monitors regulators, triages changes and drafts obligations for your team. That is this article, and the discipline it automates is regulatory horizon scanning and monitoring.

2. Compliance of agents. Governing the AI systems your company deploys: inventories, oversight, the EU AI Act. This is a different problem for a different buyer; start with our EU AI Act checklist for deployers.

3. The "AI compliance officer." Mostly a human job title (the officer who governs AI), sometimes a fantasy (an AI that is the officer). The fantasy version fails on accountability grounds, covered below with citations.

This article is about intent 1. Unlike a chatbot with a compliance system prompt (for why a general chatbot is the wrong tool for the watching, see can you use ChatGPT for regulatory horizon scanning?), an agent runs unprompted on a trigger, works through a multi-step plan, uses tools, writes durable records, and can fail in ways that are logged. Full taxonomy: agentic AI in compliance.

Every component ships with a limitation column, because a matrix without one is an ad

The matrix has one row per component of the RegWatch monitoring workflow, split into two tables. It covers what is in the shipped product today and nothing else. The limitation column comes from how we know these components fail.

Copy the tables into your own evaluation document, keep the eight columns, empty the rows, and score every vendor against them.

What each component does

Component Trigger and cadence Inputs it reads What it actually does What it writes
Company profile agent: turns your website and a brief into a structured company profile User action at onboarding Company URL, onboarding brief Reads public company information; drafts business lines, jurisdictions and regulatory topics; proposes watchlist scope A draft company profile for you to review
Source suggestions (a purpose-built pipeline, not an agent): finds the official channels worth watching Onboarding, and when you add jurisdictions or topics to a watchlist Watchlist topics and jurisdictions Searches for official publication channels and returns candidates whose jurisdiction coverage is validated Suggested sources for you to accept or reject
Monitoring agent: retrieves what regulators published, inside a strict date window The watchlist's schedule (daily, twice a week, weekly or monthly), or a manual run The watchlist's topics, jurisdictions and sources; a date window (today plus a lookback floor) Reads official listing pages first, follows in-window items to the item page or PDF, and supplements with topic and local-language searches where the watchlist allows broader search; records what it could not reach Findings: title, summary, verbatim excerpt, published date, source URL and a proposed urgency
Triage agent: scores each finding against your profile and writes reasoning in plain language Follows every monitoring run The finding, your company profile, custom exclusions, urgency rules and similar existing alerts Scores relevance from 0 to 100, sets urgency, decides accept, reject or defer, links duplicates, and writes reasoning for compliance officers, not engineers An alert with "Why this matters", business impact and recommended actions; or a suppressed finding and no alert
Compliance Assistant: answers questions from your own records User prompt (chat) The organization's alerts, obligations and policy text, up to three attachments, and web search Answers with structured citations back to those records, a retrieval-augmented pattern Chat answers with citations

How each component is controlled and audited

Component Human checkpoint Known limitations and failure modes What gets recorded
Company profile agent You review and edit the profile in the onboarding wizard before monitoring depends on it Only as good as your public web presence; thin sites produce generic profiles; conglomerates can be over-scoped The profile you approved
Source suggestions You choose which suggested sources to activate Paywalled and login-gated regulators, portal-only publication, and feeds that rot; suggestions are candidates, not a coverage guarantee Health and last successful check for each monitored source
Monitoring agent Findings are candidates: nothing reaches your inbox until triage, and people decide from there Republished pages carrying fresh dates can mis-date items (the reason for a second date check when results are saved); source-side errors and truncated responses; very large source lists can exhaust a run's budget; strict source mode is a best-effort breadth pass, not a guarantee that every publication was seen Findings with source URL, dates and verbatim excerpt; source health; failed runs, counted on the dashboard
Triage agent People disposition every alert: acknowledge, assign, dismiss with a reason or convert to an obligation Over-broad custom exclusions cause over-suppression; relevance judgment inherits the quality of the profile On each alert: the score, urgency, reasoning, applied exclusions and related alerts
Compliance Assistant A person reads the answer; consequential decisions stay with the officer Bounded by what your team has captured in the platform; answers can be wrong; not a legal-advice engine The exchange, with the citations it used

Every row ends in a person. That is the human-in-the-loop column, and it is where the liability model lives: agents propose, and the obligations register binds only after a human decision.

A monitoring run record should let an auditor check the work a year later

Whichever vendor you evaluate, a complete record of an agent run should hold these fields, and the vendor should be able to show you one.

Field What it holds What it lets you check
Status A fixed vocabulary, for example queued, running, succeeded, partial, failed, cancelled or budget_exceeded Failure is a queryable state, not an absence of output
Start and finish times Timestamps for the run Whether the run finished inside its window
Agent and model Which component ran, on which model Which system made the decision
Tokens and cost Input, cached input and output tokens, and the cost in US dollars What a run costs, and whether the pricing survives your source list
Steps An ordered list, each with a kind (such as planning, model call, tool call, retrieval or output), a duration, token counts and the input and output The sequence of work, in order
The planning step The pinned inputs, including the run date and the earliest publication date that counts as new Date-window enforcement is visible, not asserted
Tool calls Tool name, arguments, result, status (such as ok, error or timeout) and duration Every search and fetch, and every source that failed
The output step The structured batch of findings, with an outcome for each source Negative results are recorded, not skipped silently
Error The verbatim error payload The failure in the provider's or the source's own words

Three things a CCO should check in that record.

Date-window enforcement is visible, not asserted. The planning step pins the date window the run was judged against, so when an agent calls an item "new," the record shows exactly what "new" meant. In RegWatch, a second date check when results are saved rejects anything outside the window.

Negative results are recorded and checked. A source that returns nothing should not be a silent skip, because silence is a coverage hole you discover during an examination, while a logged skip is a to-do you can act on. In RegWatch, the server accepts a "nothing relevant" outcome only when it is backed by at least two distinct searches and notes on what was listed or blocked; if the accounting is incomplete, the run is marked partial and that source's schedule does not advance.

The economics are legible. A run record should carry tokens and cost, and a per-run budget should stop a run that grows out of proportion to its scope. If your vendor cannot tell you what a run costs, they do not know either.

Triage writes its reasoning down for the accepts and records the suppressions

The monitoring agent's findings are candidates. The triage agent decides what a human actually sees, and the record of that decision is the direct answer to the hallucination and defensibility objection.

An accepted finding becomes an alert with a relevance score, an urgency, a plain-language "Why this matters" written against your company profile, the business impact, recommended actions and the exclusions that applied. That is an argument for relevance to this specific firm, the one an officer would otherwise write by hand and an examiner reads when asking "why did you act on this?"

Here is the shape of a good note, written for this article for a fictional EU payment institution. It is an illustration, not an export from a customer account:

Commission Delegated Regulation (EU) 2025/532 applies from 22 July 2025, with no transitional provision for existing contracts, though changes to them must be made in a timely manner. Your profile lists card acquiring and processing run on ICT providers that subcontract, so the contract-content and notification duties apply to you. Recommended next step: a clause-gap review of the in-scope ICT contracts and an update to the register of information. Relevance high; urgency high.

The other half is a suppression. When triage rejects a finding, the finding is marked suppressed and no alert is created, so nobody's inbox carries it. Picture a national planning-policy consultation that mentions payment methods in passing, surfacing on the watchlist of a fictional payments company. The monitoring agent surfaced it in good faith, because the keywords match. Triage suppresses it: planning policy is environment intelligence, not a compliance obligation for a payments institution. A suppression is a decision too, so ask any vendor, us included, how it answers "what did you decide not to act on, and why?" for an item like this one. Most manual monitoring programs cannot answer that question at all.

How much of what an agent surfaces becomes an alert depends on your scope and your profile. If nearly everything is suppressed, the scope is too loose; if almost nothing is, the profile is probably too thin. No vendor can tell you that number in advance, and any vendor quoting an unsourced accuracy percentage is quoting marketing. Measure it on your own watchlist in a parallel run.

Failure is a status, not a silence

In a well-built run record, failure is a queryable state with a verbatim error payload. Take the example every large source list eventually produces: a run over a very large register of sources across many jurisdictions that keeps searching until it hits its budget. Three things matter to a CCO. The run is stopped by a control, not by luck: a per-run budget that keeps a sprawling register from becoming an unbounded bill. It emits nothing, so no partial, unverified findings leak into the pipeline. And it is re-runnable, at a narrower scope or a deliberately raised cap, decided by a person looking at the record.

The rest of the failure taxonomy is just as unglamorous. Sources return server errors or truncated responses. The model provider rate-limits a request, or returns a refusal, which RegWatch treats as a failed turn and never as a completed one. A run that cannot account for every source finishes as partial, and failed runs are counted on the dashboard's watchlist health panel. When infrastructure breaks, the evidence of how much monitoring did not happen should be preserved. Ask your current provider for their equivalent. An answer of "our monitoring always works" means they do not log failures.

This is also my answer to the consensus statistics in this space. Oliver Wyman's February 2026 report says its benchmarks and client engagements show AI-enabled compliance functions can reduce time spent on manual, repetitive tasks by 50% to 70% and improve risk detection by up to four times, without disclosing a sample or a method. Maybe, but an automation percentage with no failure log is unfalsifiable. The operative question is whether you can see what the agent did, what it cost and what it failed to do.

No agent replaces the compliance officer, because accountability is not delegable

The query "AI compliance officer" mostly means a human who governs AI. The latent question, whether the agent can be the officer, is real, and the answer is no on three grounds you can cite.

Accountability regimes name humans. The FCA has declined to write AI-specific rules, relying instead on existing frameworks, including the Consumer Duty and the Senior Managers and Certification Regime, where identified senior humans carry personal regulatory accountability (the FCA's AI approach page, published in September 2025 alongside its AI Live Testing feedback statement, says it does not plan to introduce extra regulations for AI). An agent cannot hold an SMF function. In the EU, AI Act Article 14 requires high-risk AI systems to be designed so that natural persons can oversee them effectively.

The error rates are not officer-grade. Stanford's peer-reviewed benchmark of purpose-built legal AI research tools found hallucinated answers on 17% (Lexis+ AI) to 33% (Westlaw AI-Assisted Research) of queries as tested in 2024. That is lower than the rates Stanford cited for general-purpose models on legal queries, but still "1 in 6 or more" (Magesh et al. and the Stanford HAI summary; the tools have shipped updates since). Domain tuning shrinks hallucination; it does not retire it, and the engineering controls that reduce each error class are covered in LLM accuracy on regulatory text.

The precedent for unsupervised AI output is a sanction. In Mata v. Avianca (S.D.N.Y., decided 22 June 2023), lawyers filed non-existent, ChatGPT-generated case citations and were jointly sanctioned $5,000 (opinion). The court noted there is nothing inherently improper about using a reliable AI tool for assistance; its problem was that the lawyers abandoned their responsibility to verify the output and then stood by it. That is the architecture the checkpoint column exists to prevent.

What agents compress is the layer below judgment. In the 2023 Cost of Compliance report from Thomson Reuters Regulatory Intelligence, 62% of respondents said they spend between 1 and 7 hours in an average week tracking and analyzing regulatory developments (report), and CUBE's 2025 survey of 2,000+ senior leaders found that 74% take more than a year to implement new regulations. That reading, triage and drafting layer is what the matrix describes. The judgment layer (is this material, do we act, who owns it) stays with the officer, working from written reasoning instead of raw feeds.

Regulators are not blocking compliance agents; they are testing them with firms

The permission question comes up in every buying conversation. The evidence, dated:

  • Adoption is mainstream. The Bank of England and FCA's joint survey (published 21 November 2024) found that 75% of UK financial services firms already use AI, with another 10% planning to within three years, and 95% adoption in insurance and 94% at international banks (BoE and FCA survey). Compliance Week's latest Inside the Mind of the CCO survey reported that nearly all responding compliance professionals use AI in their day-to-day tasks, although many said their organizations were not adequately addressing its risks (Compliance Week, 10 June 2026).
  • The FCA is testing with firms, not writing bans. The FCA says it does not plan to introduce extra regulations for AI and will rely on existing frameworks. It published feedback statement FS25/5 on AI Live Testing on 9 September 2025 and started working with the first cohort in October 2025. Second-cohort applications ran from 19 January to 24 March 2026, and testing with that cohort began in April 2026. A regulator running live AI trials with supervised firms is not hostile to compliance automation.
  • The EU AI Act does not list internal compliance monitoring as high-risk. Under Regulation (EU) 2024/1689, Annex III lists eight areas: biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration and border control, and justice and democratic processes. A tool that monitors regulatory publications for a compliance team is not in any of them. Employment is the edge case: point 4(b) covers AI used to monitor and evaluate the performance and behavior of workers, so a tool that scores staff conduct could fall inside it. Scope your deployment description carefully. The high-risk dates themselves moved: the Annex III rules now apply from 2 December 2027 under Regulation (EU) 2026/1744, in force since 27 July 2026. The Article 4 AI-literacy duty has applied since 2 February 2025 and survived that amendment in softer form: providers and deployers must take measures to support their staff's AI literacy, without guaranteeing any specific level.
  • Usage for regulatory change is still thin. KPMG's July 2025 analysis, citing an earlier edition of the Compliance Week survey, notes that 31% of respondents used AI to improve policies and procedures, while a further 13% used it to keep pace with regulatory changes (KPMG). The constraint is trustworthy operationalization, which is an evidence problem, and run records are that evidence.

Our survey of how compliance teams track regulatory changes compares today's methods: spreadsheets, feeds, consultants and agents.

Six database questions expose an agent you cannot audit, ours included

Ask them in order, and do not accept a slide as an answer to a database question:

  1. "Show me one complete run record." Steps, timestamps, tokens, cost. If the record does not exist, the "agent" is a scheduled script with a chat interface.
  2. "What is your run status vocabulary? Is failure a state or a silence?" You want to hear something like succeeded, failed, cancelled and budget_exceeded, and to see a real failed run with its error payload.
  3. "Where does a human approve before an obligation is created?" Map the answer onto the checkpoint column. No checkpoint means their liability model is your liability.
  4. "What did your system suppress for a customer like me last week, and why?" Explainable filtering is the product; a vendor that only shows accepts is showing half the system.
  5. "What does a run cost, in tokens and dollars?" Per-run economics predict whether the pricing survives your source list.
  6. "What happens when a source changes its page structure or goes behind a login?" Honest answers involve health monitoring, error rates and logged skips, not "our coverage is comprehensive."

In RegWatch, two answers are on screen for every customer: Convert to obligation is a step a person triggers, with decisions written to a tamper-evident, append-only audit log (question 3), and each monitored source shows its health and last successful check (question 6). Ask us the other four in a demo, as you would any vendor. The tools comparison applies this rubric across the field. The questions are vendor-neutral on purpose: a competitor who answers all six with real records has earned the shortlist spot.

Every statement in this article about how RegWatch works describes a field, a status or a checkpoint in the product.

This article is general information, not legal advice.

Questions

What do AI agents actually do in regulatory compliance?

They run a pipeline: build a company profile, suggest and validate regulatory sources, retrieve newly published items on a schedule inside strict date windows, score each item against the business with written reasoning, and hand accepted changes to people as alerts they can turn into tracked obligations. An assistant answers questions grounded in the firm's own records, with citations. People keep the consequential decisions: dismissing an alert with a reason, assigning it and creating the obligation.

Can an AI agent replace a compliance officer?

No. Accountability is not delegable: the FCA relies on existing regimes such as the Senior Managers and Certification Regime rather than AI-specific rules, and EU AI Act Article 14 requires effective human oversight of high-risk systems. Purpose-built legal AI tools still hallucinated on 17% to 33% of benchmark queries in Stanford's 2024 testing. Agents compress reading, triage and drafting; the officer owns judgment and sign-off.

How accurate are AI agents at detecting regulatory changes?

I know of no public benchmark for regulatory-change agents, so treat any unsourced accuracy percentage as marketing. Ask for evidence you can inspect instead: a complete run record, a logged failure with its error, what the system suppressed for a customer like you and why, and what a run costs. Failure modes worth naming are mis-dated republications, unreadable sources and budget-capped runs.

Do regulators allow compliance teams to use AI agents?

Yes, with accountability attached. The Bank of England and FCA found in a November 2024 survey that 75% of UK financial services firms already use AI. The FCA declined to write AI-specific rules and runs AI Live Testing with firms. Under the EU AI Act, a tool that monitors regulatory publications for a compliance team is not in an Annex III high-risk category, though employment-related deployments can be.

What is the difference between a compliance AI agent and a chatbot or copilot?

Trigger, persistence and evidence. A chatbot answers when prompted and leaves a transcript. A compliance agent runs on a schedule inside strict date windows, writes durable records such as findings, alerts and obligations, and leaves a run record with its steps, token counts, cost and a status vocabulary in which failure is a state. If you cannot audit the output a year later, it was a conversation, not an agent.

Terms in this guide

Sources

  1. FCA, Our approach to AI (first published 8 September 2025) accessed 30 Sep 2026
  2. FCA FS25/5, AI Live Testing feedback statement (9 September 2025) accessed 30 Sep 2026
  3. FCA announces second cohort of AI Live Testing accessed 30 Sep 2026
  4. Bank of England and FCA, Artificial intelligence in UK financial services 2024 accessed 30 Sep 2026
  5. Regulation (EU) 2024/1689 (EU AI Act), EUR-Lex accessed 30 Sep 2026
  6. European Commission, Regulatory framework for AI accessed 30 Sep 2026
  7. Stanford HAI: AI on trial, legal models hallucinate in 1 out of 6 (or more) benchmarking queries accessed 30 Sep 2026
  8. Magesh et al., Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv:2405.20362; Journal of Empirical Legal Studies 2025) accessed 30 Sep 2026
  9. Mata v. Avianca, Inc., No. 1:22-cv-01461 (S.D.N.Y. 22 June 2023), opinion and order on sanctions accessed 30 Sep 2026
  10. Oliver Wyman, Reimagining compliance with agentic AI (February 2026) accessed 30 Sep 2026
  11. KPMG, How AI is poised to reshape compliance functions (July 2025), citing Compliance Week, Inside the Mind of the CCO 2025 accessed 30 Sep 2026
  12. Compliance Week, Inside the Mind of the CCO survey finds most compliance teams allow AI use, but governance lags behind (10 June 2026) accessed 30 Sep 2026
  13. Regulation (EU) 2026/1744 (Digital Omnibus on AI), EUR-Lex accessed 30 Sep 2026
  14. CUBE, The Cost of Compliance Report 2025 accessed 30 Sep 2026
  15. Thomson Reuters Regulatory Intelligence, 2023 Cost of Compliance report (PDF) accessed 30 Sep 2026

See which of this month’s changes apply to you.

Book a session on the regulators and markets you name.

Book a demo