How do you measure AI hallucinations?

Define the task and reference evidence, label the claims in each answer, and report errors with explicit denominators. Track unanswered requests and action outcomes separately so a low error rate does not hide missing coverage.

Updated October 11, 2026

What exactly are you measuring?

Choose a property before choosing a score. FActScore divides generated text into atomic facts and measures the percentage supported by a reliable knowledge source. Its reported human evaluation concerns biographies; adopting the general idea for business workflows requires your own evidence and labeling rules. Read the paper abstract.

For this guide, use supported, contradicted and unresolved as distinct claim labels. These are our proposed reporting labels, not an assertion that every benchmark uses them. State whether a missing source counts as a failure, an unresolved result, or an exclusion; do not silently remove it from the denominator.

Which denominators should you report?

An answer may combine supported and unsupported claims; that is the motivation for the atomic-fact approach described in FActScore. Answer quality and responsiveness also differ across systems in the legal-tool study. FActScore · Legal-tool study.

What does a worked example look like?

Illustrative numbers, not a benchmark: suppose 100 requests produce 80 substantive answers and 20 abstentions. Of the 80 answers, 8 contain at least one error under your rubric. The answer error rate is 8 ÷ 80 = 10%, and the abstention rate is 20 ÷ 100 = 20%.

Now suppose those answers contain 200 evaluated factual claims: 180 supported, 12 contradicted and 8 unresolved. Claim support is 90%, contradiction is 6%, and unresolved claims are 4%. Do not describe this as 90% whole-workflow accuracy: the unit is a claim.

Report zero denominators as “not applicable,” not a perfect score. Keep a refusal that includes factual assertions in the claim review as well as recording its response status.

How should you build the test set?

Our recommended set includes common tasks, missing information, false premises, policy changes, conflicting documents and examples with consequential mistakes. Record the model, prompt, retrieval settings, source versions and date. Keep a held-out set separate from examples used to tune the system.

Have reviewers use a written rubric and record disagreements. For automated judging, compare a sample with human judgments before relying on the aggregate. Document the evaluator, exclusions and uncertainty; do not present its decisions as unquestionable ground truth.

Can you compare rates across studies?

Only make a direct comparison after checking task, models, dates, source access, labeling rules and denominators. FActScore’s biography evaluation and the legal-tool evaluation address different settings. Neither is a current universal model ranking. FActScore · Legal-tool study.

Use the prevention guide to turn observed errors into control tests. Start with the system-boundary definition if the claim being tested is unclear.

Reading scope

Frequently asked questions

Is an unsupported claim necessarily false?

Not necessarily. In this proposed rubric, distinguish contradiction from inability to establish support. Report both instead of treating missing evidence as a proved falsehood.

Does a test with zero errors prove a system is hallucination-free?

Report zero observed errors for the stated test set, versions and rubric. The observation alone does not establish behavior on every other input.

Sources

  1. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Association for Computational Linguistics (2023-12)
  2. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, arXiv (2024-05-30)