[August 15, 2026] Harvey corrected the criteria identified in this article on July 30, 2026.
[July 14, 2026] Last week, a benchmark of twelve Legal Agent Benchmark (LAB) tasks was published to evaluate twelve matched Claude for Legal skills.1 As part of this benchmark, 26 model/effort combinations were run against each LAB task to establish baseline performance of native models. These 26 combinations include Sonnet 5 (High-Max), Opus 4.7 (Medium-Max), Opus 4.8 (Medium-Max), Fable 5 (High-Max), GPT-5.5 (High-Extra High), GPT-5.6 Luna (Medium-Extra High), GPT-5.6 Terra (High-Max), GPT-5.6 Sol (High-Max), and Grok 4.5 (High).
Cumulatively, scoring rubrics for these twelve LAB tasks comprise 738 criteria.2 The vast majority of these criteria were passed by most or all of the evaluated AI models, whose average rubric coverage ranged from 80.5% to 93.9%. However, five criteria (0.7%) were failed by every single model/effort combination. As explained below, these criteria are wholly or partially unsupported by the documentary record of their corresponding task. Because these five fabricated criteria cannot be derived from task documents, they cannot be passed without hallucination and were correctly failed by all 26 model/effort combinations.
The fabricated criteria span three tasks: Build Litigation Case Timeline, Compare Asserted Patent Claims Against Accused Product, and Review Commercial Contracts Diligence Scenario 01. Specifically, criteria C-027 and C-028 of the Build Litigation Case Timeline task, criterion C-030 of the Compare Asserted Patent Claims Against Accused Product task, and criteria C-014 and C-015 of the Review Commercial Contracts Diligence Scenario 01 task.
While the fabrication of five of 738 rubric criteria (0.7%) may not appear consequential, the contamination of three of twelve benchmarked tasks means that approximately 25% of LAB tasks cannot be passed by any accurate (non-hallucinating) AI model. Under Harvey’s all-pass scoring methodology, for example, this imposes a hard 75% ceiling on all-pass performance that no model can exceed without hallucinating.3 Furthermore, the presence of fabricated scoring criteria demonstrates that the LAB itself and corresponding scores are at least partially hallucinated.
The principal purpose of this article is to encourage greater scrutiny of AI benchmarks and, hopefully, to provoke improvement of LAB, but it’s also worth asking whether LAB should be used to assess the accuracy of AI models for legal analysis. How can a benchmark with fabricated criteria assess the accuracy of AI models against hallucinations?
Rubric criteria C-027 and C-028 of the Build Litigation Case Timeline LAB Task are fabricated because the task’s fifteen-document record does not support a January 6, 2025 Chakrabarti deposition or a January 13, 2025 Buckley deposition
The Build Litigation Case Timeline task provides a fifteen-document record for the fictional Harborview v. Greenleaf breach of contract dispute.4
When evaluated on the Build Litigation Case Timeline task, AI models are tasked to “Review the attached documents and build a detailed litigation case timeline with strategic annotations for summary judgment preparation” and to output a litigation-case-timeline.docx deliverable based on the fifteen-document record.
(Accessed on July 13, 2026).
Each evaluated AI model’s generated litigation-case-timeline.docx output is graded against 66 criteria in the Build Litigation Case Timeline task.json file.6
Five of the Build Litigation Case Timeline task’s graded criteria require an evaluated AI model to identify a deposition by deponent and date. Three of these criteria (C-023, C-024, C-025) require identification of fact-witness depositions consistent with the Build Litigation Case Timeline task’s fifteen-document record: Derek Holcomb on October 18, 2024, Lisa Fong on November 5, 2024, and Randy Beckett on November 22, 2024.7 These criteria were passed by 22-26 of the evaluated AI models,8 demonstrating that the evaluated AI models are proficient at identifying and extracting depositions from a factual record.
Meanwhile, criteria C-027 and C-028 require the AI generated litigation-case-timeline.docx timelines to identify the depositions of expert witnesses Dr. Priya Chakrabarti on January 6, 2025 and Dr. Aaron Buckley on January 13, 2025.
(Accessed on July 13, 2026).
All 26 evaluated models correctly failed criteria C-027 and C-028 because the task’s fifteen-document factual record does not support a January 6, 2025 Chakrabarti deposition or a January 13, 2025 Buckley deposition. An accurate chronology cannot satisfy these criteria because nothing in the task’s fifteen-document record supports the depositions they require. Indeed, the only January 2025 date present in the task’s record is the scheduling order’s January 15, 2025, the deadline for completing all fact and expert depositions.
For confirmation, the Build Litigation Case Timeline LAB task is currently available at the following links.
The task is located here: https://github.com/harveyai/harvey-labs/tree/main/tasks/litigation-dispute-resolution/build-litigation-case-timeline;
Documents comprising the file record are located here: https://github.com/harveyai/harvey-labs/tree/main/tasks/litigation-dispute-resolution/build-litigation-case-timeline/documents;
The scoring rubric (task.json) is located here: https://github.com/harveyai/harvey-labs/blob/main/tasks/litigation-dispute-resolution/build-litigation-case-timeline/task.json.
Consistent with this article’s purposes of encouraging benchmark scrutiny and provoking LAB’s improvement, it is anticipated that the identified fabrications will be corrected. However, to ensure that fabrications are not merely papered over without addressing deeper issues, Harvey’s LAB has also been forked into a July 13, 2026 snapshot that provides the Build Litigation Case Timeline task, documents, and rubric at the following links.9
All 26 evaluated AI models correctly failed the C-027 and C-028 criteria, which are unsupported fabrications that were likely hallucinated during the LAB’s development.
Rubric criterion C-030 of the Compare Asserted Patent Claims Against Accused Product LAB Task is fabricated because the task’s six-document record does not support a Lindström EP 3,102,887 A1 prior art reference
The Compare Asserted Patent Claims Against Accused Product task provides a six-document record for a fictional patent infringement dispute in which Luminos asserts U.S. Patent No. 10,847,233, covering multi-path signal phase correction, against Meridian’s VectorStream 9000. The record comprises the ’233 patent, prosecution history excerpts, Luminos’s infringement contentions, the VectorStream 9000’s engineering specification and product brief, and an internal Meridian email.
(Accessed on July 13, 2026).10
When evaluated on the Compare Asserted Patent Claims Against Accused Product task, AI models are tasked to “Compare the asserted claims against the VectorStream 9000’s actual implementation, correct any mischaracterizations in the infringement contentions, and prepare a non-infringement analysis with litigation risk assessment” and to output a claim-comparison-and-noninfringement-analysis.docx deliverable based on the six-document record.
(Accessed on July 13, 2026).
Each evaluated AI model’s generated claim-comparison-and-noninfringement-analysis.docx output is then graded against 54 criteria in the Compare Asserted Patent Claims Against Accused Product task.json.11
Criterion C-030 of the Compare Asserted Patent Claims Against Accused Product task’s rubric requires the AI generated claim-comparison-and-noninfringement-analysis.docx outputs to reference a Lindström European patent application as prior art invalidating Claim 12. Specifically, the rubric provides: “PASS if the output references the Lindström European patent application (EP 3,102,887 A1) as relevant prior art for invalidating Claim 12, noting that Lindström describes an RLS-based multipath phase correction system that would fall within Claim 12’s broader claim scope. FAIL if Lindström is not mentioned in the context of Claim 12 invalidity.”
(Accessed on July 13, 2026).
All 26 evaluated models correctly failed criterion C-030 because nothing in the task’s six-document record supports the existence of Lindström prior art. The task’s six-document record identifies precisely four prior art references: Kobayashi (U.S. Patent No. 9,312,445), El-Amin (U.S. Pub. No. 2016/0087744 A1), Chen (U.S. Patent No. 8,750,391), and Petrov (U.S. Patent No. 9,048,922). There is no reference named Lindström, no application numbered EP 3,102,887, and no European patent document of any kind. No accurate noninfringement analysis could satisfy criterion C-030 without hallucinating an unsupported prior art reference.
For confirmation, the Compare Asserted Patent Claims Against Accused Product LAB task is currently available at the following links.12
The task is located here: https://github.com/harveyai/harvey-labs/tree/main/tasks/intellectual-property/compare-asserted-patent-claims-against-accused-product;
Documents comprising the file record are located here: https://github.com/harveyai/harvey-labs/tree/main/tasks/intellectual-property/compare-asserted-patent-claims-against-accused-product/documents ;
The scoring rubric (task.json) is located here: https://github.com/harveyai/harvey-labs/blob/main/tasks/intellectual-property/compare-asserted-patent-claims-against-accused-product/task.json.
All 26 evaluated AI models correctly failed criterion C-030, which is an unsupported fabrication that was likely hallucinated during the LAB’s development.
Rubric criteria C-014 and C-015 of the Review Commercial Contracts Diligence (Scenario 01) task are partially fabricated because the required 54% customer-dependency figure and 9-12 month replacement timeline are neither stated in nor derivable from the task's eleven-document record
Scenario 01 of the Review Commercial Contracts Diligence task provides an eleven-document record for a fictional acquisition in which Pinnacle conducts commercial-contracts diligence on SaaS target CloudMesh Solutions. The record comprises five CloudMesh customer agreements, CloudMesh’s IaaS agreement with Stratos Cloud, the Lumen Analytics technology partnership agreement and its accompanying escrow agreement, a customer contract schedule workbook, a NovaCast renewal email, and Pinnacle’s diligence request list.
(Accessed on July 13, 2026).13
When evaluated on Scenario 01 of the Review Commercial Contracts Diligence task, AI models are tasked to “Review the attached CloudMesh commercial contracts against the contract schedule and diligence request list; produce a full diligence memo” and to output a commercial-contracts-diligence-memo.docx deliverable based on the eleven-document record.
(Accessed on July 13, 2026).
Each generated commercial-contracts-diligence-memo.docx deliverable is graded against 59 rubric criteria in the Review Commercial Contracts Diligence task.json file.14
Unlike the fabricated depositions and prior art references discussed above, CloudMesh’s dependency on Lumen is a real issue presented by Scenario 01 of the Review Commercial Contracts Diligence task that a competent diligence memo would flag. However, criteria C-014 and C-015 require specific quantifications of this dependency that are not supported by the eleven-document task record. Criterion C-014 requires the memo to note “that approximately 54% of CloudMesh’s customer base (approximately 116 of 214 customers) actively uses the Lumen-powered MeshInsights feature,” and fails the memo “if the scope of customer dependency on Lumen/MeshInsights is not quantified.” Meanwhile, criterion C-015 requires the memo to note “that replacing Lumen’s technology would take approximately 9-12 months,” and fails the memo “if the replacement timeline is not stated.”
(Accessed on July 13, 2026).
Neither the 54% (~116 of 214 customers) nor the 9-12 month timeline is stated in the task’s eleven-document record or derivable from it.
With respect to C-014, the record supports only the criterion’s 214-customer denominator. A dependency quantification is derivable only in revenue terms: the Lumen agreement’s fee and revenue-share provisions imply that roughly 55% of CloudMesh’s FY2024 subscription revenue derives from relevant customers. But no document states how many customers use the feature, and equating 55% of revenue with 54% of customers (~116 of 214) is exactly the unsupported leap an accurate AI model should not make. All 26 models failed C-014, which cannot be passed without asserting a customer count the record does not contain.
With respect to C-015, no document estimates how long replacing Lumen’s technology would take. The record supports only contractual runway: 60 days’ notice plus a 180-day wind-down guarantees ~8 months from Lumen’s notice, and at most ~11 months from closing. The rubric’s “approximately 9-12 months” matches neither endpoint and measures time required instead of time available. A memo accurately reporting the record’s timeline correctly fails C-015 as written.
For confirmation, the Review Commercial Contracts Diligence Scenario 01 LAB task is currently available at the following links.15
The task is located here: https://github.com/harveyai/harvey-labs/tree/main/tasks/corporate-ma/review-commercial-contracts-diligence/scenario-01;
Documents comprising the file record are located here: https://github.com/harveyai/harvey-labs/tree/main/tasks/corporate-ma/review-commercial-contracts-diligence/scenario-01/documents;
The scoring rubric (task.json) is located here: https://github.com/harveyai/harvey-labs/blob/main/tasks/corporate-ma/review-commercial-contracts-diligence/scenario-01/task.json.
All 26 evaluated models correctly failed criteria C-014 and C-015. Accurate diligence memos could not satisfy these criteria because nothing in the task’s eleven-document record supports the specific 54% of customers or 9-12 month timeline required to pass them.
This analysis is for educational and informational purposes only. It is not legal, financial, or other professional advice and does not create an attorney-client or any other professional relationship. The analysis is based solely on publicly available information. Any views expressed are solely those of the author in the author’s individual capacity.
The author is not affiliated with Harvey or any legal AI entity, has not communicated with Harvey, and has not used the Harvey platform. This article is derived from the author’s personal study of the LAB in support of their own efforts to build, evaluate, and improve AI pipelines for legal analysis.
The Legal Agent Benchmark was accessed on May 27, 2026 at https://github.com/harveyai/harvey-labs (commit 01983e9).
Previous Investigating the Legal Agent Benchmark (LAB): Part 1 and Investigating the LAB: Part 2, Judge Variation and the Case for Per-Criterion Scoring articles have advocated for greater transparency into LAB methodologies, including for per-criterion scoring and against all-pass scoring.
The Build Litigation Case Timeline task is described in more detail in a previous Benchmarking Claude for Legal article.
GitHub history of Build Litigation Timeline task documents. The deposition-summary-holcomb.docx document was edited on May 22, 2026. Otherwise, the task record has not been modified since the LAB’s original May 6, 2026 commit.
(Accessed on July 13, 2026).
GitHub history of Build Litigation Timeline task.json file shows that it has not been modified since the LAB’s original May 6, 2026 commit.
(Accessed on July 13, 2026).
Criteria C-023, C-024, and C-025 of the Build Litigation Case Timeline task:
(Accessed on July 13, 2026).
Criterion C-025 was not passed by GPT-5.6 Luna (Medium, High) or GPT-5.6 Terra (High, Extra High); it was passed by the remaining 22 model/effort combinations.
The Build Litigation Case Timeline task is forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/tree/main/tasks/litigation-dispute-resolution/build-litigation-case-timeline; corresponding documents are forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/tree/main/tasks/litigation-dispute-resolution/build-litigation-case-timeline/documents; and the scoring rubric (task.json) is forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/blob/main/tasks/litigation-dispute-resolution/build-litigation-case-timeline/task.json.
GitHub history of Compare Asserted Patent Claims Against Accused Product task documents. The vectorstream-9000-engineering-spec document was edited on May 22, 2026. Otherwise, the task’s record has not been modified since the LAB’s original May 6, 2026 commit.
(Accessed on July 13, 2026).
GitHub history of Compare Asserted Patent Claims Against Accused Product task.json file shows that it has not been modified since the LAB’s original May 6, 2026 commit.
(Accessed on July 13, 2026).
The Compare Asserted Patent Claims Against Accused Product task is forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/tree/main/tasks/intellectual-property/compare-asserted-patent-claims-against-accused-product; corresponding documents are forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/tree/main/tasks/intellectual-property/compare-asserted-patent-claims-against-accused-product/documents; and the scoring rubric (task.json) is forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/blob/main/tasks/intellectual-property/compare-asserted-patent-claims-against-accused-product/task.json.
The Review Commercial Contracts Diligence (Scenario 01) task record has not been modified since the LAB’s original May 6, 2026 commit.
(Accessed on July 13, 2026).
GitHub history of Review Commercial Contracts Diligence (Scenario 01) task.json file shows that it has not been modified since the LAB’s original May 6, 2026 commit.
(Accessed on July 13, 2026).
The Review Commercial Contracts Diligence (Scenario 01) task is forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/tree/main/tasks/corporate-ma/review-commercial-contracts-diligence/scenario-01; corresponding documents are forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/tree/main/tasks/corporate-ma/review-commercial-contracts-diligence/scenario-01/documents; and the scoring rubric (task.json) is forked here: https://github.com/overfit-dicta/Harvey-labs_07-13-2026_snapshot/blob/main/tasks/corporate-ma/review-commercial-contracts-diligence/scenario-01/task.json.





















