An asterisk can qualify a slogan. It cannot run an evidence check. If a supplier promises hallucination-free work, I want the system boundary in the first paragraph of the specification.

My argument is simple: precise guarantees are useful. Unspecified guarantees are difficult to test. The difference is the contract between a claim and the evidence that could defeat it.

What is the strongest case for the claim?

There is a serious engineering argument for bounded guarantees. AWS describes Automated Reasoning checks that validate responses against user-defined policies, identify supporting rules and variable assignments, and detect contradictions. That is a specific property with a defined reference. It deserves to be evaluated on those terms. AWS documentation.

The strongest counterargument to my skepticism is that well-defined tasks and explicit rules can make a meaningful assurance possible. I agree. A buyer should welcome a precise assurance and examine how the implementation enforces it.

Where does the promise become too large?

The trouble starts when a check against one policy becomes shorthand for the truth of everything in an answer. AWS’s retrieved documentation describes policy-relative validation. It also says the checks do not protect against prompt injection and validate the supplied content as-is. Those are boundaries to include in the review, not reasons to dismiss the technique. AWS documentation.

Likewise, a factuality measure has a particular target. FActScore breaks generated text into atomic facts and measures support against a reliable knowledge source. Its abstract describes biography evaluations. A good score on that task does not, by itself, tell a buyer whether an invoice workflow authorizes the right payment. That last distinction is an inference about scope, not a finding from a payment experiment. FActScore.

What would I put in the contract?

I would ask the supplier to complete four sentences:

  1. The task is… Identify the decisions and outputs the system covers.
  2. The authority is… Name the records, policy versions and conditions that count as support.
  3. The gate is… Explain which check must pass before an answer is accepted or an action is authorized.
  4. When the check cannot pass… Show the stop, clarification or human-review path.

Then demonstrate a missing record, a contradictory policy and an ambiguous input. Keep the actual output, action and exception record. These are my proposed acceptance checks; they are not a published benchmark or a certification.

What would change my mind?

A clearly scoped promise, an inspectable control and a test record that includes failures would. I do not need a grander slogan. I need to know when the workflow refuses to turn missing evidence into a completed decision.

Start with what a hallucination-free claim covers and how to measure its results. Make the boundary visible enough that another reviewer can disagree with the result.

The source record

Read the original evidence and the scope of our review.

  1. What are Automated Reasoning checks in Amazon Bedrock Guardrails?Amazon Web Services · Not stated · Accessed 2026-10-11Substantive indexed capability sections and opening limitations read; full page not retrieved.
  2. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationAssociation for Computational Linguistics · 2023-12 · Accessed 2026-10-11Indexed abstract and metadata read; full paper not read. Human evaluation described in the abstract concerns biographies.
← Latest storiesFollow the briefing ↗