What changed?
Artificial Analysis announced Harvey LAB-AA v1.1 on October 8. The update checks legal deliverables against their source documents, alongside grading by three LLM judges. A task must meet every criterion and contain no material hallucination to count toward the headline metric. Announcement.
The evaluation uses 120 private legal tasks across 24 practice areas. Agents work with instructions and case documents through Artificial Analysis’s Stirrup harness. These are benchmark tasks, not an assessment of every commercial legal product. Benchmark methodology.
Why does it matter?
Our take: a useful evaluation should make unsupported content visible even when an answer otherwise satisfies the assignment. This update makes that distinction explicit in the success criteria.
Read the score carefully: failing this combined measure does not necessarily mean an answer hallucinated. It could fail another rubric requirement. The benchmark also warns that its scores are not comparable with v1.0. Methodology.
What remains uncertain?
Private tasks limit outside inspection, and grading uses models. The retrieved methodology does not establish reliability in deployed legal practice. We read the announcement and substantive methodology passages; neither full page was available through retrieval.
For a practical review process, see how to verify AI-generated citations and how to measure hallucinations.
The source record
Read the original evidence and the scope of our review.
- Announcing Harvey LAB-AA v1.1: adding hallucination checks to raise the bar for agentic legal workArtificial Analysis · 2026-10-08 · Accessed 2026-10-11Substantive indexed article passages read; retrieved text was truncated.
- Harvey LAB-AA v1.1 Benchmark LeaderboardArtificial Analysis · Publication date not displayed · Accessed 2026-10-11Indexed background and methodology read; complete page not retrieved.