LLM Accuracy on Regulatory Text: Hallucination Risks and How to Control Them

In short

General-purpose chatbots hallucinated on 58% to 88% of verifiable questions about real federal court cases (Dahl et al., Journal of Legal Analysis, 2024), and commercial legal research tools built on retrieval still hallucinated on 17% to 33% of queries (Magesh et al., Journal of Empirical Legal Studies, 2025). Accuracy on regulatory text is an architecture property, and the checklist below has six domains and 25 controls, each with a pass/fail test.

LLMs are as accurate on regulatory and legal text as the architecture around them. The evidence forms a ladder. General-purpose chatbots hallucinated on 58% to 88% of verifiable questions about real federal court cases (Dahl, Magesh, Suzgun and Ho, Journal of Legal Analysis, 2024, testing 2023-era models). Commercial legal research tools built on retrieval-augmented generation still hallucinated on 17% to 33% of queries (Magesh et al., Journal of Empirical Legal Studies, 2025). Grounded summarization of a single supplied document does far better, at roughly 2% to 15% for most models on Vectara's September 2026 leaderboard. What changes from rung to rung is the system around the model.

So hallucination on regulatory text is an architecture problem you engineer down rather than a model property you wait out, and most vendors would rather leave that unargued. The engineering is specifiable, testable and already expected by supervisors under rules written before ChatGPT existed. The deliverable is the Hallucination-Control Checklist for Regulatory AI below: six control domains, every control with spec language for a requirements document, a pass/fail test you can run in an afternoon, and the framework it maps to.

I run RegWatch, a platform that points LLM agents at regulators, so I have a commercial interest in you believing AI can do compliance work. That is why the numbers below include the ugly ones. The oldest studies here tested 2023-era models, and no benchmark is a guarantee for any particular system, mine included. You can still test the system in front of you, and the checklist shows how.

Hallucination rates form a ladder, and architecture decides which rung you're on

Here is the evidence chain, dated, with the caveats each study's critics would raise:

Rung System architecture Measured error rate Study and date Caveats a hostile reader would raise
4 (worst) General chatbot, closed-book, asked verifiable questions about randomly sampled real federal cases 58% (GPT-4), 69% (GPT-3.5) and 88% (Llama 2) Dahl et al., Journal of Legal Analysis 16(1):64 to 93, published June 2024 The models are 2023-era and frontier models score better, but the failure mode is structural (see below)
3 Commercial legal research tools with retrieval-augmented generation 17% (Lexis+ AI and Ask Practical Law AI) to 33% (Westlaw AI-Assisted Research) Magesh et al., "Hallucination-Free?", preregistered, tested March to May 2024, published April 2025 LexisNexis noted a newer generation of its tool after testing, and Thomson Reuters said Ask Practical Law AI was not intended for primary-law research (LawSites, May 2024). The peer-reviewed finding stands: retrieval reduces hallucination and does not eliminate it
3 Commercial legal AI on 50-state statutory surveys Accuracy of 58% (Westlaw AI) and 64% (Lexis+ AI) on unemployment-insurance statute questions; the researchers' own custom pipeline reached 83% Afane et al., Stanford RegLab, February 2026 preprint The authors built the winning pipeline, and it measures accuracy on yes-or-no statute questions, not a hallucination rate. It is still the closest public evidence for multi-jurisdiction regulatory text
2 Grounded summarization of a single supplied document About 2% to 3% for the best small models and roughly 6% to 15% for most frontier models Vectara Hallucination Leaderboard, HHEM-2.3 judge, updated 22 September 2026 It measures summary faithfulness, not legal question answering. Rates for the same models rose several-fold when Vectara changed its benchmark dataset in November 2025 (GPT-5 high: 1.4% in the October 2025 archive, 15.1% now). Use the shape, grounded far below ungrounded, and never a per-model ranking
1 (target) Grounded, cited, date-windowed, human-gated workflow Residual errors exist and are caught by verification before they ship A control-stack property, not a benchmark; the checklist below is its specification Nobody publishes audited rates at this rung, which is itself the procurement question (Domain E)

Three things follow that the raw papers never spell out for compliance buyers.

First, the spread within a rung is task difficulty, not luck. A 2025 study of LLM legal explanations for influencer-marketing compliance, run with small models on 1,143 Instagram posts under the Dutch Advertising Code, reached an F1 score of up to 0.93 across the dataset, lost more than 10 points on 95 ambiguous posts, and left out legal citations in 28.57% of the explanations the researchers annotated (Gui et al., arXiv 2510.08111). AIReg-Bench generated 120 fictional but plausible excerpts of AI Act technical documentation with an LLM, had legal experts annotate which Articles each one violates, and tested whether frontier models reproduce the experts' labels; its authors offer the result as "a starting point to understand the opportunities and limitations" of such tools, not as proof they are ready. Operationally: automate the unambiguous with sampling QA, and route ambiguity to humans by design.

Second, Dahl et al. found two failures scarier than wrong answers. The models often failed to correct questions built on false legal premises, and they struggled to predict their own hallucinations, so an ungrounded model's confidence is no guide to whether it is wrong.

Third, nobody on the ladder gets to zero. The best available answer to whether the next model generation fixes this, "not by scale alone," comes from the people with the strongest incentive to say yes.

OpenAI's researchers say accuracy will never reach 100%, and the remedy is behavior, not size

In September 2025, Kalai, Nachum, Vempala and Zhang published "Why Language Models Hallucinate". They argue that hallucination arises from statistical pressures inherent in pretraining, because some facts are not learnable from the training distribution, and that it is compounded by training and evaluation procedures that "reward guessing over acknowledging uncertainty." A model that answers "I don't know" scores zero on many benchmarks, while a model that guesses sometimes scores. Their prescription is to change what gets rewarded, penalizing confident errors and giving credit for appropriate expressions of uncertainty. In OpenAI's own summary, accuracy will never reach 100% because some real-world questions are inherently unanswerable, but hallucination is not inevitable, because a model can abstain when it is uncertain.

For a Chief Risk Officer, this converts an open research question into a procurement one. If hallucination were a maturing-technology problem, the rational move would be to wait. Since the fix is behavioral, the rational move is to demand compensating controls now: grounding, so the model has something true to be faithful to; refusal behavior, so uncertainty surfaces as "not found" instead of a guess; and verification, so the residual slips get caught.

Banking supervisors have long had a frame for this. Models err, and institutions cope through validation, ongoing monitoring and effective challenge. That was the logic of SR 11-7. On 17 April 2026 the Federal Reserve, OCC and FDIC replaced it with revised interagency model risk management guidance, SR 26-2. The new text keeps the principles for traditional statistical and non-generative models, and it states that "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance," adding that a bank's own risk management and governance practices "should guide the determination of appropriate governance and controls" for tools it does not cover. That gap leaves the control specification for language models yours to write.

The failures that hurt compliance teams are misgrounded citations and wrong dates, not invented statutes

The popular image of AI failure, a chatbot inventing a "Directive (EU) 2019/2087" out of nothing, is real but also the easiest to catch, because a fabricated instrument fails the first lookup. The Magesh study documents the more dangerous type: misgrounded citations, where the cited source exists, resolves, and does not support the claim attached to it. Misgrounding survives the lazy check ("does the link work?") and fails only the expensive one ("does the source actually say this?").

Their diagnosis of why retrieval tools still fail names four causes: naive retrieval that pulls the wrong passages, retrieval of inapplicable authority (right topic, wrong jurisdiction or point in time), sycophancy toward the user's framing, and plain reasoning errors. The mix differed by tool. Naive retrieval and inapplicable authority dominated for Lexis+ AI, while reasoning errors were the largest cause for Westlaw AI-Assisted Research and Ask Practical Law AI, and sycophancy was rare. Better retrieval helps without fixing accuracy on its own, because generation still reasons badly from good sources. That is why verification and human gates belong in the design as well as in the policy.

The second underrated failure is dating. Regulators' pages often display a "last updated" or republication date rather than the original publication date, and an agent that faithfully reads the date on the page will mis-date the item. A mis-dated item does a compliance team as much damage as a hallucinated one: it fires a false "new development" alert or silently compresses an implementation window. No public benchmark we know of measures republication handling, in-force versus proposed labeling, or date-window discipline, which is why the checklist gives dates their own domain instead of treating "hallucination" as one undifferentiated risk.

The Hallucination-Control Checklist gives every control a pass/fail test and a framework hook

Spec language goes in your requirements document or vendor RFP. The test column is deliberately cheap; most take under an hour. The mapping column connects each control to a named source, because the fastest way to fund a control is to show an examiner or an auditor already assumes you have it. Frameworks referenced: NIST AI 600-1 (July 2024); FINRA Notice 24-09 (27 June 2024) and the FINRA 2026 Annual Regulatory Oversight Report (9 December 2025); the FCA's approach to AI; and the research cited above. None of them is a checklist for language models, so "maps to" means "the closest anchor," not "required by."

Domain A: Grounding and retrieval

Control What to require (spec language) How to test it (pass/fail) Maps to
Grounded generation only Factual answers are generated exclusively from retrieved source documents; parametric memory is not an acceptable source for regulatory claims Ask about a plausible but non-existent instrument (for example "Directive (EU) 2019/9999 on crypto custody"). Pass = refusal or "not found." Fail = a confident summary NIST AI 600-1 confabulation actions; FINRA Notice 24-09 on the reliability and accuracy of the AI model
Hybrid retrieval Retrieval combines keyword and semantic search so exact instrument names and article numbers are never lost to embedding fuzz Query an exact article number you know is in the corpus; pass = the retrieved passage contains that article verbatim Magesh et al. (naive retrieval as a leading cause of errors)
Grounded refusal When retrieval returns nothing relevant, the system says so; "not found" beats a guess Ask a question the corpus cannot answer. Pass = explicit no-answer with the search scope stated. Fail = an answer Kalai et al. (reward abstention); NIST AI 600-1 MEASURE
Source authority tiers Sources carry authority ranking (regulator, then official gazette, then law firm, then news) and conflicts resolve upward Seed the corpus with a blog summary that contradicts the regulator text; pass = the regulator text wins and is the citation Magesh et al. (inapplicable authority); FCA position that firms stay accountable for outcomes

Domain B: Citation and verification

Control What to require (spec language) How to test it (pass/fail) Maps to
Mandatory inline citations Every factual sentence in output carries an inline citation to a retrieved document; uncited factual output is rejected before surfacing Sample 20 outputs; any factual sentence without a citation = fail FINRA 2026 report (validation and human review of outputs); NIST AI 600-1 MEASURE
Citation resolution Citations must resolve to the actual cited document, and the passage must support the claim: this is the misgrounding check Sample 20 outputs, click every citation, read the passage against the claim, and count misgrounded cites. Pass threshold: zero Magesh et al. (misgrounded citations); Charlotin database (fabricated authority in court filings)
Quote-level pinning Direct quotes and obligation statements pin to the specific source passage, not just the document Pick 5 quoted passages; verify each exists verbatim at the pinned location NIST AI 600-1 provenance actions
Automated cite-existence check A pre-surfacing pipeline step verifies every cited document exists and is reachable before output reaches a user Ask the vendor to demonstrate the check firing on a bad citation in a live trace NIST AI 600-1 MEASURE

Domain C: Date and version discipline

Control What to require (spec language) How to test it (pass/fail) Maps to
Strict date windows Retrieval and monitoring runs receive an explicit "today" and a lookback floor; out-of-window items are rejected again at persist time (defense in depth) Ask for "what changed in the last 30 days" and verify that every returned item's original publication date falls in the window FINRA Notice 24-09 (data integrity)
Publication vs republication The system distinguishes original publication dates from republication or page-refresh dates Feed a consultation page showing a republished date; pass = the original date is surfaced No framework we reviewed names this, which is why you must
In-force vs proposed labeling Every instrument reference is labeled with status (proposed, adopted, in force, repealed) as of a stated date Query an instrument in transition (adopted, not yet applicable); pass = correct status with an "as of" date No framework we reviewed names this
"As of" stamps Every obligation statement in output carries the date at which it was true Inspect 10 obligation statements; pass = all dated Editorial norm regulators apply to their own guidance

Domain D: Human-in-the-loop gates

Control What to require (spec language) How to test it (pass/fail) Maps to
Named reviewer gate No AI output becomes an obligation, filing or client communication without a named human reviewer signing off Trace 5 recent obligations back through the workflow; pass = a named reviewer on each FINRA Rule 3110 supervision, which Notice 24-09 says applies to generative AI as to other technology
Confidence-based escalation Low-confidence or low-agreement outputs route to mandatory review rather than default acceptance Ask the vendor what happens mechanically when the model is uncertain; pass = a routing rule, fail = "the model is usually right" FINRA 2026 report (ongoing monitoring and human-in-the-loop review)
Reviewer sees the reasoning Reviewers see the model's cited reasoning and sources, not just the conclusion Sit with a reviewer; pass = they can open the reasoning and citations without leaving the review screen NIST AI 600-1 GOVERN
Sampling QA on accepted items Items accepted without edit get periodic sampled re-review, because rubber-stamping is the failure mode of every review gate Ask for the last sampling QA report; pass = it exists and has findings FINRA 2026 report ("regular checks for errors or bias")

Domain E: Vendor and procurement questions

Control What to require (spec language) How to test it (pass/fail) Maps to
Documented failure "Show me a wrong answer your system produced and how it was caught" Pass = a specific example with the catching control. Fail = "our system doesn't hallucinate," and you can end the meeting there NIST AI 600-1 MAP
Rate on your documents Hallucination rate measured on your document types, with methodology disclosed Pass = a number, a corpus description and a judge. Fail = a leaderboard screenshot NIST AI 600-1 MAP
Inspectable traces You can inspect the full agent or run trace for any output: what was retrieved, what was decided, why Request the trace for a specific output during the demo; pass = shown live FINRA 2026 report; NIST AI 600-1
Uncertainty behavior Defined, demonstrable behavior when the model is uncertain or retrieval is empty Run the Domain A non-existent-instrument test on the vendor's own demo Kalai et al.; NIST AI 600-1
Benchmark, dataset and judge named Any quoted accuracy number names the benchmark, the dataset version and the judge; published rates for the same models moved several-fold when Vectara changed its benchmark dataset Ask "which benchmark, which dataset, which judge?"; pass = a direct answer with the caveat acknowledged Basic measurement honesty; Vectara leaderboard history

Domain F: Monitoring and audit

Control What to require (spec language) How to test it (pass/fail) Maps to
Persisted reasoning artifacts Every accept, suppress or triage decision persists its reasoning and citations as a durable record Pull a 60-day-old decision; pass = the reasoning is still readable and attributable NIST AI 600-1 MANAGE
Tamper-evident audit log Agent runs land in an append-only audit log in which tampering is detectable Ask how tampering would be detected; pass = a cryptographic or equivalent mechanism, demonstrated General audit-trail practice; nothing AI-specific in the frameworks reviewed
Drift review cadence A fixed re-test set runs on a schedule (quarterly at minimum) to catch accuracy drift after model or prompt changes Ask for the last two drift reports and what changed between them NIST AI 600-1 MEASURE
Error register Caught errors feed a register that drives prompt, retrieval and routing fixes, so the loop closes Ask for the register and one fix it produced NIST AI 600-1 MEASURE and MANAGE

Twenty-five controls looks heavy, but Domains A to C are properties a system either has or lacks (test once per vendor or build), Domain D is workflow design, and E and F are procurement and operations discipline. A team can run the full test column against one system in about two days.

Regulators expect governance, not a specific control set, and none offers a safe harbor

It is a common misreading to treat FINRA's notice or the FCA's position as a "new AI regulation." Both say existing, technology-neutral rules already apply, which is more demanding in one respect and less prescriptive in another: nobody tells you which controls to run, and everybody expects you to have chosen some.

  • FINRA. Notice 24-09 (27 June 2024) says that if a firm uses generative AI tools as part of its supervisory system, its policies and procedures "should address technology governance, including model risk management, data privacy and integrity, reliability and accuracy of the AI model." It adds that it "does not create new legal or regulatory requirements or new interpretations of existing requirements." The 2026 Annual Regulatory Oversight Report (9 December 2025) defines hallucinations as "instances where the model generates information that is inaccurate or misleading, yet is presented as factual information," and describes ongoing monitoring of prompts, responses and outputs and human-in-the-loop review as practices firms may want to consider. In Regulatory Notice 26-14 (9 July 2026, a proposal to modernize Rule 2210 whose comment period closed on 11 September 2026), FINRA says generative AI communication tools can be part of a reasonably designed supervisory system "provided they are vetted, tested and monitored."
  • FCA. The AI Update (22 April 2024) says the FCA's rules and core principles do not usually mandate or prohibit specific technologies, and the FCA's AI approach page states: "We do not plan to introduce extra regulations for AI." Firms remain accountable for outcomes whatever produced them. The joint Bank of England and FCA survey (21 November 2024) found that 75% of responding firms were already using AI, with a further 10% planning to within three years.
  • US bank supervisors. As described above, the April 2026 interagency guidance (SR 26-2) leaves generative and agentic AI outside its scope and points banks to their own governance.
  • NIST AI 600-1 (July 2024) gives the risk its most useful name, confabulation, defined as "the production of confidently stated but erroneous or false content," and maps GOVERN, MAP, MEASURE and MANAGE actions to it. For vendor-neutral internal-policy language, start there.
  • EU AI Act (Regulation (EU) 2024/1689, in force 1 August 2024). Article 15 requires accuracy, robustness and cybersecurity for high-risk systems, and Article 15(3) requires the levels of accuracy and the relevant accuracy metrics to be declared in the instructions for use. After the Digital Omnibus, Regulation (EU) 2026/1744, those obligations apply from 2 December 2027 for Annex III systems and 2 August 2028 for Annex I products; the general-purpose AI obligations have applied since 2 August 2025. Be precise about scope: a regulatory-monitoring tool is typically not an Annex III high-risk system (Annex III point 8(a) covers AI used by or for judicial authorities to research and interpret facts and law), so Article 15 does not bind it directly. Use Article 15(3) as a procurement hook instead: ask any vendor to declare accuracy metrics for the intended purpose. If you deploy systems that are in scope, the deadlines are mapped in our EU AI Act compliance checklist for deployers.

And the record for uncontrolled AI output is no longer hypothetical. Damien Charlotin's AI Hallucination Cases database lists 2,097 court decisions worldwide involving hallucinated material as of 30 September 2026. In the database's export, 1,206 entries involve self-represented litigants and 825 involve lawyers, professionals trained in source verification who filed fabricated authority because the output looked checked. The database's scope also covers items other than citations, such as fabricated exhibits. Every filing with a fabricated authority would have died at Domain B, row two: click the citation, read the passage. The compliance analogue is an obligation register entry citing a provision that doesn't say what the entry claims. Nobody benchmarks that in public, so your controls have to catch it.

You can trust AI for grounded, cited, reviewed tasks, and you should not trust it for open-ended recall

Every page ranking for "can you trust AI for compliance" answers "it depends."

Trust it, at rung 2 and with Domain B and D controls, for: summarizing a supplied document; first-pass triage of retrieved items against a documented profile, with human review; comparing two versions of a text you provide; extracting structured fields (dates, thresholds, scope) from a named instrument, with citation pinning; drafting for human editing. These are grounded tasks. Raw error rates on Vectara's summarization benchmark run from about 2% to 15% for most models, far below the 58% to 88% of closed-book legal recall, and verification catches most of the residue. Note the direction of that number: it is not zero, and it moved when the benchmark got harder, so keep the sampling QA in Domain D running.

Do not trust it for: open-ended legal or regulatory research from a chat interface (rung 4, the 58% to 88% rung); any citation you have not resolved and read (the 2,097-decision rung); deadline or in-force-date extraction without Domain C discipline; ambiguous applicability judgments, where measured accuracy drops by double digits; anything uncited. The specific ways an unscoped chat interface breaks on monitoring work are documented, with a reproducible stress test, in Can you use ChatGPT for regulatory horizon scanning?.

The profession has, in aggregate, priced this about right. The Thomson Reuters Institute's 2026 AI in Professional Services Report, a survey of 1,514 professionals conducted in October and November 2025, found that four in ten respondents say their organizations are using generative AI, up from 22% the year before, while among those opposed to applying it, "worries over reliability and accuracy" are the primary reason. Practitioners believe the task fit and distrust the failure mode. Both instincts are right, and human-in-the-loop architecture reconciles them, with review effort concentrated where the evidence says errors cluster, in ambiguity, citations and dates.

Two cheap tests show which rung of the ladder your AI is actually on

Run the two cheapest tests from the checklist against whatever AI touches your compliance workflow today: the non-existent-instrument probe (Domain A) and the 20-output citation resolution audit (Domain B). Together they take half a day and tell you which rung of the ladder you are actually on, as opposed to the rung the sales deck claims. A vendor that declines to run them on a live screen has answered the question anyway.

Publishing a scorecard you do not apply to yourself is the oldest trick in vendor content, so RegWatch should be held to its own checklist. Parts of Domains B, C, D and F describe behavior you can check in the product rather than policy. The Compliance Assistant answers from an organization's own alerts, obligations and policy text, plus the attachments and web search it is given, and its answers carry structured citations (B). Monitoring runs receive an explicit "today" and a lookback floor, a second date check runs before a finding is saved, and each finding keeps its source URL, dates and a verbatim excerpt (C). Triage records every decision to accept, reject or defer, with a written "Why this matters" that a compliance officer can read and overrule, and dismissing an alert requires a reason (D). Decisions go to an append-only audit log that is hash-chained and tamper-evident (F). I would still run the two cheap tests on our own assistant, and I would like prospects to bring the Domain A and B tests to our demos, because no vendor can self-certify its own accuracy. The broader case for the agent architecture is in AI agents for regulatory compliance.

We expect models to keep improving and regulated firms to still be running every control in this checklist five years from now, because accuracy never reaches 100% and supervisory logic, from FINRA's 2026 report to NIST's confabulation profile to the FCA's insistence that existing rules apply, treats accuracy on regulatory text as a property you operate. Other pieces on AI tooling for compliance are collected in the AI compliance hub.


This article is general information, not legal advice.

Questions

How often do LLMs hallucinate on legal and regulatory questions?

The rate depends on the architecture. General-purpose chatbots hallucinated on 58% to 88% of verifiable questions about real federal court cases (Dahl et al., Journal of Legal Analysis, 2024, testing 2023-era models). Commercial legal research tools with retrieval hallucinated on 17% to 33% of queries (Magesh et al., Journal of Empirical Legal Studies, 2025). Grounded summarization of one supplied document is far better, roughly 2% to 15% for most models on Vectara's September 2026 leaderboard.

Does retrieval-augmented generation (RAG) stop hallucinations?

It reduces them but does not eliminate them. Stanford's preregistered benchmark of commercial legal research tools found hallucination on 17% to 33% of queries, traced to naive retrieval, inapplicable authority, sycophancy and reasoning errors. Reasoning errors were the largest cause for two of the three tools. That residual risk is why citation verification and human-in-the-loop gates exist.

Will hallucinations go away as models improve?

Not by scale alone. OpenAI researchers argued in September 2025 that models hallucinate partly because training and evaluation reward guessing over admitting uncertainty. They say accuracy will never reach 100% and that the fix is to penalize confident errors and reward appropriate abstention. In practice that means grounding, refusal behavior and verification, not waiting for a bigger model.

Do regulators allow compliance teams to use LLMs?

Yes, under existing technology-neutral rules; none of the regulators cited here bans them. FINRA says its rules apply to generative AI as to any technology, and the FCA says it does not plan extra AI-specific regulation. In the US, the April 2026 interagency model risk guidance leaves generative AI out of scope and points banks to their own governance. None of them offers a safe harbor for unsupervised reliance.

What should we demand from an AI compliance-monitoring vendor?

Five things: a wrong answer their system produced and how it was caught; a hallucination rate measured on your document types with the methodology, dataset and judge disclosed; inspectable per-decision reasoning and citations; defined refusal behavior when the system is uncertain; and human-review gates plus an audit trail. These follow from NIST AI 600-1's confabulation controls and long-standing model validation practice.

Terms in this guide

Sources

  1. Dahl, Magesh, Suzgun and Ho, Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Journal of Legal Analysis 16(1) (2024) accessed 30 Sep 2026
  2. Dahl et al., Large Legal Fictions (arXiv 2401.01301, version 2 abstract) accessed 30 Sep 2026
  3. Magesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies 22(2) (2025) accessed 30 Sep 2026
  4. Stanford HAI: AI on Trial, legal models hallucinate in 1 out of 6 or more benchmarking queries accessed 30 Sep 2026
  5. LawSites: Stanford will augment its study finding that AI legal research tools hallucinate in 17% of queries, as some raise questions about the results (28 May 2024) accessed 30 Sep 2026
  6. Afane, Hariri, Ouyang and Ho, Benchmarking Legal RAG: The Promise and Limits of AI Statutory Surveys (arXiv 2603.03300) accessed 30 Sep 2026
  7. Vectara Hallucination Leaderboard (HHEM-2.3, updated 22 September 2026) accessed 30 Sep 2026
  8. Vectara Hallucination Leaderboard, archived October 2025 version (branch hhem-2.3-old-dataset) accessed 30 Sep 2026
  9. Vectara: Introducing the next generation of Vectara's Hallucination Leaderboard (19 November 2025) accessed 30 Sep 2026
  10. Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate (arXiv 2509.04664) accessed 30 Sep 2026
  11. OpenAI: Why language models hallucinate (September 2025) accessed 1 Oct 2026
  12. Gui et al., Evaluating LLM-Generated Legal Explanations for Regulatory Compliance in Social Media Influencer Marketing (arXiv 2510.08111) accessed 30 Sep 2026
  13. Marino et al., AIReg-Bench: Benchmarking Language Models That Assess AI Regulation Compliance (arXiv 2510.01474) accessed 30 Sep 2026
  14. FINRA Regulatory Notice 24-09 (27 June 2024) accessed 30 Sep 2026
  15. FINRA 2026 Annual Regulatory Oversight Report, GenAI section (9 December 2025) accessed 30 Sep 2026
  16. FINRA Regulatory Notice 26-14: proposed changes to modernize Rule 2210 (9 July 2026; comment period closed 11 September 2026) accessed 30 Sep 2026
  17. FCA: AI Update (22 April 2024) accessed 30 Sep 2026
  18. FCA: AI and the FCA, our approach accessed 30 Sep 2026
  19. Bank of England and FCA: Artificial intelligence in UK financial services 2024 (21 November 2024) accessed 30 Sep 2026
  20. NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative AI Profile (July 2024) accessed 30 Sep 2026
  21. Federal Reserve SR 26-2: Revised Guidance on Model Risk Management (17 April 2026), with the interagency attachment accessed 30 Sep 2026
  22. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Articles 15 and 113, and Annex III accessed 30 Sep 2026
  23. Regulation (EU) 2026/1744 of 8 July 2026 (Digital Omnibus on AI), published 24 July 2026, in force 27 July 2026 accessed 30 Sep 2026
  24. Damien Charlotin: AI Hallucination Cases database (2,097 cases, last updated 30 September 2026) accessed 1 Oct 2026
  25. Damien Charlotin: AI Hallucination Cases database, CSV export accessed 1 Oct 2026
  26. Thomson Reuters Institute: 2026 AI in Professional Services Report accessed 30 Sep 2026

See which of this month’s changes apply to you.

Book a session on the regulators and markets you name.

Book a demo