Take an answer apart before giving it one score. That is the useful starting point in FActScore: a long response can mix supported and unsupported facts, so a single all-or-nothing label loses information. The paper’s abstract defines a finer unit of evaluation. FActScore.
This explainer is based on the accessible abstract and metadata. It explains the metric’s stated meaning, not the full implementation or a reproduction of the experiments.
What is the unit being measured?
An atomic fact is an individual factual assertion extracted from a generated response. FActScore asks whether a reliable knowledge source supports each one, then computes the percentage supported. The abstract describes human evaluation of generated biographies and an automated estimator combining retrieval with a language model. FActScore.
The denominator is therefore the facts being evaluated in the response. It is not automatically the set of facts a user needed to accomplish a task. That distinction follows from the definition of the metric.
Why does the denominator matter?
Consider a deliberately simplified illustration. A task requires five facts. An answer includes two, both supported. Under a bare supported-facts fraction, those two facts receive full support credit; the three omitted requirements remain a separate completeness question.
This is a teaching example, not a published result or a claim about how every FActScore implementation handles length or abstention. We did not inspect the full paper or code for those details.
| Review question | Evidence to inspect |
|---|---|
| Are the included claims supported? | Each claim and its reference passage |
| Does the answer cover the requested task? | Required items and which were answered |
| Did the system decline to answer? | Response disposition and stated reason |
| Did a business action succeed? | Authorization and destination-system record |
These are complementary review fields we recommend. A single fraction cannot substitute for all four.
Does the setting change the interpretation?
Yes. The FActScore abstract’s human evaluation concerns biographies. A separate preregistered legal-research study evaluates professional legal tools and reports differences in accuracy and responsiveness. The settings and questions differ; neither abstract supplies a universal reliability score for business automation. FActScore · Legal-research study.
We do not reproduce either study’s model scores here. The retrieved abstracts are enough to explain why task and metric definitions matter, but not enough to adjudicate every implementation choice or assess today’s products.
How should a team use this insight?
Before testing, write down the task, authoritative sources, unit of judgment and aggregation rule. Retain claim-level decisions so reviewers can inspect disagreements. Report coverage and abstention alongside support, with explicit denominators.
If reviewers cannot locate the reference passage, record the support check as unresolved. Do not silently convert missing evidence into a positive label. Our measurement guide explains the review sequence, and the citation guide covers source identity and status.
The source record
Read the original evidence and the scope of our review.
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationAssociation for Computational Linguistics · 2023-12 · Accessed 2026-10-11Indexed abstract and metadata read; full paper not read. Human evaluation described in the abstract concerns biographies.
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsarXiv · 2024-05-30 · Accessed 2026-10-11Indexed abstract and metadata read; no full protocol or current product assessment.
