What is an AI hallucination?
In this reference, an AI hallucination is generated content presented as factual or source-based that is false or unsupported by the evidence required for the task.
How should the term be used?
This is our working definition. State the evaluation rule whenever you use it: false factual content and a claim unsupported by a supplied document are related but different findings.
FActScore evaluates support for individual factual claims; the legal-tool reliability paper studies hallucinations in a specific professional setting. Their scopes illustrate why a definition must travel with the metric. FActScore · Legal-tool study.
What is a simple example?
In a hypothetical policy summary, an assistant says “the reimbursement limit is $500” when the provided policy contains no such limit. That statement is unsupported by the supplied source. Determining the actual limit requires the appropriate authority.
Do not use the label to explain every software bug, missing step or unauthorized action. Record the observable failure and the evidence for it.
Read how to measure hallucinations and what a hallucination-free claim should cover.
Reading scope
- Association for Computational Linguistics: Abstract and bibliographic metadata read; the reported human evaluation concerns biographies.
- arXiv: Abstract and metadata read; findings refer to the tools evaluated in this study.
Sources
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Association for Computational Linguistics (2023-12)
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, arXiv (2024-05-30)