What is the update?
Vectara’s Hallucination Leaderboard displays a September 22, 2026 update. Its introduction describes evaluating hallucinations when models summarize a supplied document. The visible table includes hallucination rate, factual consistency, answer rate and average summary length. Leaderboard.
That is one particular reliability question: does the summary stay faithful to its source?
What does a different benchmark measure?
Artificial Analysis describes AA-Omniscience as a factual-knowledge and calibration evaluation. Its scoring rewards accuracy, penalizes bad guesses and rewards abstention when uncertain. It is asking a different question from whether a summary faithfully reflects an attached document. AA-Omniscience.
Our take: before sharing a model’s “hallucination rate,” attach the task, evaluator and answer policy. Compare results within a consistent evaluation, and examine how often the model answered as well as how often it was wrong.
What can we conclude?
These sources describe two evaluation designs. They do not supply one interchangeable scale for all business AI. Vectara reports results using its own evaluator; our review did not independently reproduce them.
We recovered the leaderboard’s introduction and part of its table, not its full methodology. We therefore do not publish model rankings or infer its sample size and refusal handling. AA-Omniscience’s retrieved page supplied substantive methodology summaries, not its complete paper.
Use our measurement guide to define the task and denominator before interpreting a score.
The source record
Read the original evidence and the scope of our review.
- Hallucination LeaderboardVectara · 2026-09-22 · Accessed 2026-10-11Displayed update date, introduction and part of the results table read; full methodology unavailable.
- AA-Omniscience: Knowledge and Hallucination BenchmarkArtificial Analysis · Publication date not displayed · Accessed 2026-10-11Indexed background and methodology summary read; full research paper not read.