AI hallucination
An AI hallucination is output that a language model states fluently and confidently but that is false or not supported by its sources.
An AI hallucination is output from a language model that is stated fluently and confidently but is false, or is not supported by the sources the model was given. NIST's generative AI profile uses the term confabulation, defined as "the production of confidently stated but erroneous or false content" (NIST AI 600-1, July 2024), and notes that "hallucinations" and "fabrications" are the colloquial names.
Measured rates depend on the architecture around the model. On verifiable questions about real US federal court cases, general-purpose chatbots were wrong 58% to 88% of the time (Dahl, Magesh, Suzgun and Ho, 2024, testing 2023-era models). Commercial legal research tools that retrieve sources before answering still hallucinated on 17% to 33% of queries (Magesh et al., 2025). Grounding a model in retrieved text, the idea behind retrieval-augmented generation, reduces the problem without removing it. OpenAI researchers have argued that training and evaluation reward guessing over acknowledging uncertainty (Kalai et al., 2025), which is why behavior such as abstaining when unsure matters more than model size.
For compliance work, invented statutes are the easy case, because they fail the first lookup. The harder failures are a real citation attached to a passage that does not say what the output claims, and a wrong date, such as a republication date read as the original publication date. The controls follow from those failure modes: require a citation for every claim and resolve it to the passage, keep verbatim excerpts, check dates against a defined window, accept "not found" as a valid answer, and route low-confidence or consequential outputs to a human reviewer. The LLM accuracy article sets these out as a checklist with a pass or fail test for each control.
Sources
- NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (July 2024) accessed 30 Sep 2026
- Dahl, Magesh, Suzgun and Ho, Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Journal of Legal Analysis 16(1) (2024) accessed 30 Sep 2026
- Magesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies 22(2) (2025) accessed 30 Sep 2026
- Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate (arXiv 2509.04664, September 2025) accessed 30 Sep 2026
