How do you measure AI hallucinations?
Define the task and reference evidence, label the claims in each answer, and report errors with explicit denominators. Track unanswered requests and action outcomes separately so a low error rate does not hide missing coverage.
What exactly are you measuring?
Choose a property before choosing a score. FActScore divides generated text into atomic facts and measures the percentage supported by a reliable knowledge source. Its reported human evaluation concerns biographies; adopting the general idea for business workflows requires your own evidence and labeling rules. Read the paper abstract.
For this guide, use supported, contradicted and unresolved as distinct claim labels. These are our proposed reporting labels, not an assertion that every benchmark uses them. State whether a missing source counts as a failure, an unresolved result, or an exclusion; do not silently remove it from the denominator.
Which denominators should you report?
- Claim support. Proposed calculation: Supported claims ÷ all evaluated factual claims. What it helps you see: How much of the answer has evidence.
- Answer error rate. Proposed calculation: Substantive answers containing at least one defined error ÷ all evaluated substantive answers. What it helps you see: How often an answer includes a failure.
- Abstention rate. Proposed calculation: Requests declined without a substantive answer ÷ all requests. What it helps you see: How often the system does not answer.
- Unresolved claim rate. Proposed calculation: Claims that cannot be adjudicated ÷ all evaluated factual claims. What it helps you see: Gaps in evidence or evaluation.
- Action outcome. Proposed calculation: Confirmed, failed and unresolved actions, with counts. What it helps you see: Whether reported work matches the workflow record.
An answer may combine supported and unsupported claims; that is the motivation for the atomic-fact approach described in FActScore. Answer quality and responsiveness also differ across systems in the legal-tool study. FActScore · Legal-tool study.
What does a worked example look like?
Illustrative numbers, not a benchmark: suppose 100 requests produce 80 substantive answers and 20 abstentions. Of the 80 answers, 8 contain at least one error under your rubric. The answer error rate is 8 ÷ 80 = 10%, and the abstention rate is 20 ÷ 100 = 20%.
Now suppose those answers contain 200 evaluated factual claims: 180 supported, 12 contradicted and 8 unresolved. Claim support is 90%, contradiction is 6%, and unresolved claims are 4%. Do not describe this as 90% whole-workflow accuracy: the unit is a claim.
Report zero denominators as “not applicable,” not a perfect score. Keep a refusal that includes factual assertions in the claim review as well as recording its response status.
How should you build the test set?
Our recommended set includes common tasks, missing information, false premises, policy changes, conflicting documents and examples with consequential mistakes. Record the model, prompt, retrieval settings, source versions and date. Keep a held-out set separate from examples used to tune the system.
Have reviewers use a written rubric and record disagreements. For automated judging, compare a sample with human judgments before relying on the aggregate. Document the evaluator, exclusions and uncertainty; do not present its decisions as unquestionable ground truth.
Can you compare rates across studies?
Only make a direct comparison after checking task, models, dates, source access, labeling rules and denominators. FActScore’s biography evaluation and the legal-tool evaluation address different settings. Neither is a current universal model ranking. FActScore · Legal-tool study.
Use the prevention guide to turn observed errors into control tests. Start with the system-boundary definition if the claim being tested is unclear.
Reading scope
- Association for Computational Linguistics: Abstract and bibliographic metadata read; the reported human evaluation concerns biographies.
- arXiv: Abstract and metadata read; findings refer to the tools evaluated in this study.
Frequently asked questions
Is an unsupported claim necessarily false?
Not necessarily. In this proposed rubric, distinguish contradiction from inability to establish support. Report both instead of treating missing evidence as a proved falsehood.
Does a test with zero errors prove a system is hallucination-free?
Report zero observed errors for the stated test set, versions and rubric. The observation alone does not establish behavior on every other input.
Sources
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Association for Computational Linguistics (2023-12)
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, arXiv (2024-05-30)