Can You Use ChatGPT for Regulatory Horizon Scanning? What Breaks and Why
In short
In a peer-reviewed study of legal hallucination, public LLMs got 58% to 88% of verifiable questions about real federal court cases wrong (Dahl et al., Journal of Legal Analysis, 2024), and ChatGPT cannot run regulatory horizon scanning because coverage, dating and an audit trail are properties of a monitoring system, not of a chat model. It remains a good tool for explaining a regulation you name.
You can use ChatGPT to explain a regulation you already know exists. You cannot use it to run regulatory horizon scanning, the scheduled, documented sweep of everything your regulators publish. That gap is measurable. In a peer-reviewed study of legal hallucination, public LLMs got 58% to 88% of verifiable questions about real federal court cases wrong (Dahl, Magesh, Suzgun and Ho, Journal of Legal Analysis, 2024). And when I re-ran an access check against four regulator websites on 30 September 2026, the failure started earlier than most people assume: at retrieval, before the model gets a chance to hallucinate.
I run a company that points LLM agents at regulators every day, so I see where they break, including ours, which run on the same kind of base models and survive production only because of engineering controls. Each failure mode below carries a primary source.
ChatGPT is good at explaining a document you name
Early advice on this topic, written in 2023, assumed a chat window with no live web access. As of October 2026, ChatGPT can search the live web, open sites in a built-in browser, run scheduled tasks that monitor for changes, keep work in projects and read uploaded files, and OpenAI's own help page cautions that search results and citations can be incomplete, outdated or incorrect. Ask it to explain DORA's register of information, compare a consultation draft against the final rule you pasted in, or turn a 40-page instrument into a briefing outline, and it performs well, often better than the junior analyst who would otherwise do the first pass.
So the steelman is real: for comprehension of a document you hand it or name precisely, ChatGPT is a legitimate tool. What it is not built to do is the job compliance teams actually mean by horizon scanning: every source you answer to, on schedule, correctly dated, with a record of what was seen and dismissed. That is a coverage and dating task with an evidence trail attached, and judged that way, it breaks in five specific places.
Horizon scanning breaks in five documented places
| Failure mode | What it looks like | Root cause | Documented evidence | The control that catches it |
|---|---|---|---|---|
| Retrieval gaps | The regulation it never saw; the answer may omit whatever couldn't be fetched without saying so | Regulator sites throttle automated access: blocked search endpoints, rate thresholds, bot checks, PDF-only formats | Access check of the FCA, EUR-Lex, federalregister.gov and SEC.gov, 30 September 2026 (below) | A curated source inventory with per-source fetch monitoring, so a failed fetch is an incident, not silence |
| Dating errors | "New" items that are republications; deadlines read off the wrong date | No date discipline at generation; republication pages and PDFs degrade dates | Regulator pages that show a republication date instead of the original (below) | Strict date windows (today plus a lookback floor) enforced again at persist time |
| Coverage illusion | A "scan" sampled from three or four famous sources, presented as complete | Search-engine mediation; no defined source universe | The Federal Register alone ran 106,109 pages and 3,248 final rules in 2024, and 60,917 pages and 2,441 final rules in 2025 (CEI, from Office of the Federal Register data) | A written source register; coverage measured against it |
| Fabricated citations | Confident references to instruments, cases and paragraphs that do not exist | Models are rewarded for guessing under uncertainty (OpenAI's own researchers) | 58% to 88% (Dahl et al., 2024); 17% to 33% even with retrieval (Magesh et al.); Mata; Ayinde | Output grounded only in retrieved documents, with mandatory citations; uncited output rejected |
| No scan record | Same question, different answers; the chat log shows the conversation, not the sources checked, the failed fetches or the dismissals | Each run generates a fresh answer, and the log records the conversation rather than the scan | noyb's complaint against OpenAI, which says OpenAI declined to correct a false birth date (2024); BoE and FCA: 46% of respondent firms report only a partial understanding of the AI they use (2024) | Recorded reasoning for every decision; a record of which sources were read and which failed; a tamper-evident audit log |
Retrieval fails first: regulator websites are hostile terrain for live fetching
Here is what four regulator front doors said to automated access when I checked on 30 September 2026. Robots.txt files and bot defenses change without notice, so treat every line as dated:
- FCA (robots.txt): disallows every query-string URL (
Disallow: /*?) in the default rules that apply to crawlers in general (User-agent: *), and its/search/path as well. That pattern covers search results, the exact surface a live lookup needs. - EUR-Lex: a plain scripted request, including one for its robots.txt file, was answered with a JavaScript challenge (an HTTP 202 response with no content). A client that cannot run the challenge gets nothing.
- federalregister.gov (robots.txt): disallows its document, regulation, article and public-inspection search paths, and a plain scripted request for a document page was redirected (HTTP 302) to an unblock page.
- SEC.gov: a plain scripted request with no declared user agent returned a 403 "Request Rate Threshold Exceeded" page, which says automated access must comply with the SEC's privacy and security policy and points to its fair access guidance. A request that identified itself was served normally.
None of the three robots.txt files I could read names GPTBot, OAI-SearchBot or ChatGPT-User, so it would be wrong to claim regulators "block ChatGPT." OpenAI documents four user agents, including GPTBot for training, OAI-SearchBot for search results and ChatGPT-User for actions a user triggers in ChatGPT, and it says robots.txt rules may not apply to that last one. What an agent meets at regulator sites is friction, not a wall. The consequence is quieter and worse than blocking: a human researcher who cannot reach a document knows they missed it, while a chat answer can present whatever survived the fetch as if it were the scan.
The coverage illusion compounds from there. The Federal Register closed 2024 at 106,109 pages with 3,248 final rules, then fell to 60,917 pages and 2,441 final rules in 2025, and that is one publication stream in one country. An answer assembled from a search engine's view of a few prominent sources is a sample pretending to be a census. A monitoring program inverts the order: define the source universe first, then measure coverage against it. Operating one teaches that fetch failures are routine events to detect. In our own product, a source that stops answering is flagged in its health status, so a gap appears as a fact on the record instead of a silence.
Dating errors are the failure mode nobody benchmarks
Hallucination gets the headlines; dates cause the quieter misses. Regulators re-publish. A guidance page gets a template migration and suddenly shows last week's date; a consultation is corrected and re-issued. An agent that reads the page date is then wrong about what is new. The model that reads the page is the same kind of model ChatGPT runs on, so the fix has to sit outside the model. Every monitoring run in our product receives an explicit today and a lookback floor, and a second check when results are saved rejects out-of-window items the model let through. It is defense in depth, built on the assumption that the model will get dates wrong.
A general-purpose chat session enforces no equivalent window. Ask it what a regulator published in the last 30 days and the answer can blend three populations without telling them apart: genuinely new items, republished pages carrying fresh timestamps, and recollections from training data that browsing never verified. Browsing also masks a training cutoff that is months old: when a live fetch fails or is skipped, the model can answer from training data, and the answer does not always make clear which statements came from a fetched page and which from memory. For horizon scanning, an item dated wrong is as bad as an item missed: it fires a false alarm or quietly eats your implementation runway.
Fabricated citations are structural, not a bug being patched out
Dahl and colleagues found hallucination rates of 69% for ChatGPT-3.5 and 88% for Llama 2 on specific, verifiable questions about randomly selected federal court cases, with GPT-4 at 58%, and the models often failed to push back on questions built on false legal premises. More sobering for anyone who thinks grounding fixes this: Stanford's preregistered benchmark of purpose-built, retrieval-augmented legal research tools found that Lexis+ AI hallucinated on 17% of queries and Westlaw's AI-Assisted Research on 33% (Magesh et al., arXiv:2405.20362, peer-reviewed in the Journal of Empirical Legal Studies, 2025; Stanford's plain-language summary rounds the result to "1 out of 6 or more"). Both vendors questioned aspects of the study in 2024, and without a persisted trail of what was retrieved and why, even vendors and researchers cannot settle what the tools actually did.
The consequences are in the case law now. In Mata v. Avianca (S.D.N.Y., 22 June 2023), the court imposed a $5,000 penalty, jointly and severally on the two lawyers and their firm. The court noted there was nothing inherently improper in using a reliable AI tool for assistance; the sanction was for bad faith, because the lawyers submitted non-existent opinions with fake quotes and then kept standing by them after the court questioned whether they existed. In Ayinde v Haringey; Al-Haroun v QNB ([2025] EWHC 1383 (Admin), 6 June 2025), five of the authorities cited in one case did not exist, and in the other 18 of 45 citations were fictitious. The Divisional Court said the threshold for contempt proceedings was met for one barrister but decided not to start them, referred the lawyers to their regulators, and warned that general-purpose AI tools are not capable of conducting reliable legal research. Damien Charlotin's AI Hallucination Cases database now lists more than 2,000 such decisions worldwide (2,097 when last updated on 30 September 2026).
This will not be patched out of general chat models. Researchers at OpenAI and Georgia Tech argue that hallucination is structural, because language models are optimized to be good test-takers and guessing when uncertain improves test performance. And OpenAI's own Terms of Use say the quiet part in contract language: you should not rely on output as a sole source of truth or factual information, or as a substitute for professional advice, and you must evaluate output for accuracy, using human review as appropriate. The terms for Europe say the same in their own version, and OpenAI's business terms make the customer "solely responsible" for evaluating the accuracy and appropriateness of output. The engineering patterns that actually control the risk are the subject of LLM accuracy on regulatory text.
A scan you cannot replay is not a monitoring program
Run the same prompt in two fresh sessions and you will often get two different scans: different sources, different items, sometimes different dates. The chat history keeps both conversations, and on Enterprise plans OpenAI's Compliance API can export them, but a conversation log does not record which sources were in scope, which fetches failed or what was dismissed and why. For a compliance function the deliverable is the defensible record, and a chat log records the conversation, not the scan. The sharpest illustration is noyb's GDPR complaint (April 2024): ChatGPT repeatedly gave a wrong birth date for a public figure, and according to the complaint, OpenAI refused the request to correct or erase it, saying it was not possible to correct the data and that factual accuracy in large language models remains an area of active research. As of 1 October 2026, noyb's case tracker lists the complaint as pending; Ireland's Data Protection Commission has handled it since January 2025.
Supervisors already see the gap. The Bank of England and FCA's 2024 survey (November 2024) found that 75% of responding firms already use AI, and that 46% report only a "partial understanding" of the AI they use, largely because it arrives through third-party models. An unauditable scanning process run on a tool you only partly understand is the pattern that number describes.
Run the stress test before you take my word for it
Every claim above is checkable in twenty minutes. Pick a regulator you know cold and run these five prompts against ChatGPT, or any tool, including a vendor demo and including ours:
- "List everything [regulator] published in the last 30 days." Tests coverage and dating. Check against the regulator's own publications page.
- "Give the official citation and URL for each item." Click every link and verify every citation exists; anything invented is fabrication caught in the open.
- Open a fresh session and run prompt 1 again, verbatim. Tests reproducibility. Diff the two lists.
- "What did [regulator] publish on [a specific date you already know]?" Tests recall versus retrieval: does it fetch, or remember, and can you tell?
- "Show me the exact paragraph of [instrument] that imposes [an obligation you know well]." Tests whether it can ground that answer in the primary text.
Grade pass or fail per failure mode. Anything missing or mis-dated in 1 and 4 is a retrieval or dating failure; any dead link or invented citation in 2 and 5 is fabrication; any material difference in 3 means the scan cannot be reproduced for an examiner. Expect a plain chat session to stumble hardest on prompts 2 and 3. If a vendor will not sit through this protocol on a live screen, that is also a result.
Keep ChatGPT for comprehension; never assign it the watching
Keep ChatGPT for what it is good at: explaining instruments you name, first-draft summaries, structured comparisons of documents you supply, always with a human verifying every citation, in line with OpenAI's terms, which require you to evaluate output for accuracy.
For the watching, the job spec falls straight out of the failure modes, and it is an engineering spec, not a model choice: a defined source inventory; scheduled runs inside strict date windows; per-item triage that records plain-language reasoning; and a loop where every accepted item becomes an obligation with an owner, a deadline and evidence. Run that spec on a spreadsheet and a disciplined team, or on agents. The honest comparison of both paths is in AI agents for regulatory compliance and the tools roundup.
On my own product, I will stick to claims I can back with how it behaves rather than with positioning. RegWatch's agents share the base-model weaknesses described in this essay. The difference is that RegWatch is built as a system of record for regulatory change, with several of the controls this essay argues for built into the workflow. Every run has a date window, and an independent date check runs before results are saved. Findings are stored with their source URL, dates and a verbatim excerpt. Triage writes a reasoned "Why this matters" that explains which changes apply to your business and why, and a compliance officer can read and overrule it. Owners, obligations, evidence and decisions are recorded, and each decision goes to a tamper-evident, append-only audit log. The compliance assistant answers with citations back to your own alerts and obligations. None of that makes a model infallible, but it makes the failures visible and the decisions reviewable.
My bet is that chat assistants will keep getting better at explaining regulation, and that in three years they will still not be doing horizon scanning. Coverage, dating discipline and auditability are properties of a monitoring system, and a chat interface that keeps no monitoring record, enforces no source universe and is rewarded for guessing when uncertain is the structural opposite of one. If you disagree, the stress test settles it either way. We keep related analysis of this class of tooling in the AI compliance hub.
This article is general information, not legal advice.
Questions
Can ChatGPT automatically monitor regulatory changes?
Only partly. As of October 2026, ChatGPT searches the web automatically or on request, and scheduled tasks can rerun a prompt on a timetable and monitor for changes. But its records are built around the conversation: chat history and, on Enterprise, OpenAI's Compliance API log what was asked and answered, not which sources were in scope, which fetches failed or which items were dismissed and why, so a scheduled prompt falls short of a monitoring program.
How often does ChatGPT make up legal or regulatory citations?
A peer-reviewed study found that public LLMs hallucinated on 58% to 88% of verifiable questions about real federal court cases (Dahl et al., Journal of Legal Analysis, 2024). Even purpose-built legal research tools that use retrieval produced incorrect information 17% to 33% of the time in Stanford's benchmark, although both vendors questioned aspects of the study in 2024.
Is it against the rules to use ChatGPT for compliance work?
No regulator bans it outright. But OpenAI's Terms of Use say you should not rely on output as a sole source of truth or as a substitute for professional advice, and courts have sanctioned unverified AI citations: $5,000 jointly in Mata v. Avianca (2023), and in Ayinde (EWHC, June 2025) the court said the contempt threshold was met for one barrister, though it started no proceedings. Use it with verification and never as the record.
Can ChatGPT read the actual text of regulations on EUR-Lex or the Federal Register?
Yes, for openly served pages, within limits. As of 30 September 2026, EUR-Lex answers plain scripted requests with a JavaScript challenge, federalregister.gov redirects them to an unblock page, SEC.gov returns a rate-threshold error to clients that do not identify themselves, and the FCA disallows query-string URLs in robots.txt. A browser that runs JavaScript can get past some of this; what a live fetch cannot reach, the answer may omit without saying so.
What does a regulatory monitoring platform do that ChatGPT doesn't?
Four things: a defined source inventory, scheduled runs inside strict date windows, per-item triage that records reasoning a human can inspect, and a loop that turns each accepted change into an owned obligation with a deadline and evidence. All four are engineering controls, not model capabilities.
Terms in this guide
Sources
- Dahl, Magesh, Suzgun and Ho, Large Legal Fictions, Journal of Legal Analysis 16(1), 2024 accessed 30 Sep 2026
- Stanford HAI: AI on trial, legal models hallucinate in 1 out of 6 (or more) benchmarking queries accessed 30 Sep 2026
- Magesh et al., Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv:2405.20362) accessed 30 Sep 2026
- Mata v. Avianca, Inc., No. 1:22-cv-01461 (S.D.N.Y. 22 June 2023), opinion and order on sanctions accessed 30 Sep 2026
- Ayinde v London Borough of Haringey and Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin) accessed 30 Sep 2026
- Damien Charlotin, AI Hallucination Cases database (2,097 cases, last updated 30 September 2026) accessed 1 Oct 2026
- OpenAI, Why Language Models Hallucinate (arXiv:2509.04664) accessed 30 Sep 2026
- OpenAI Terms of Use (rest of world, effective 1 January 2026) accessed 30 Sep 2026
- OpenAI Europe Terms of Use (updated 16 January 2026) accessed 30 Sep 2026
- OpenAI Services Agreement (business terms, effective 1 January 2026), section 4.3 accessed 1 Oct 2026
- OpenAI crawler documentation accessed 30 Sep 2026
- OpenAI Help Center, Searching the web with ChatGPT accessed 1 Oct 2026
- OpenAI Help Center, Scheduled tasks in ChatGPT accessed 1 Oct 2026
- ChatGPT documentation: Browser accessed 1 Oct 2026
- ChatGPT documentation: Projects accessed 1 Oct 2026
- ChatGPT documentation: Compliance API (Enterprise) accessed 1 Oct 2026
- noyb: ChatGPT provides false information about people, and OpenAI can't correct it accessed 30 Sep 2026
- noyb: case C078, OpenAI OpCo, LLC (case status page) accessed 1 Oct 2026
- Bank of England and FCA, Artificial intelligence in UK financial services 2024 accessed 30 Sep 2026
- FCA robots.txt accessed 30 Sep 2026
- Federal Register robots.txt accessed 30 Sep 2026
- SEC.gov developer resources and fair access guidance accessed 30 Sep 2026
- CEI, Ten Thousand Commandments 2026 (Federal Register data for 2025) accessed 30 Sep 2026
- CEI, Ten Thousand Commandments 2025 (Federal Register data for 2024) accessed 30 Sep 2026
