In May, Harvey released an open-source (MIT licensed) Legal Agent Benchmark (LAB)1 for evaluating AI models on legal work. In that release, Harvey invited researchers and lawyers "to help test the benchmark, pressure-test the tasks, and tell us where the signal is clearest or missing." This article accepts that invitation by beginning an independent investigation2 in support of the LAB mission "to evaluate and improve agent capabilities for supporting legal work."
LAB is available on GitHub and includes myriad tasks organized into practice-specific folders such as “litigation-dispute-resolution,” “corporate-ma,” “data-privacy-cybersecurity,” and “corporate-governance.” Each task is provided as a data file (task.json) containing an instruction that prompts the evaluated models to perform the task, one or more requested outputs (e.g., a diligence-summary-memo) to capture the AI work product, and a rubric of criteria for judging the output. Most tasks also include a folder of documents simulating the factual context (e.g., a case file or data room) on which the task is performed.
A previous article used the “build-litigation-case-timeline” task from the “litigation-dispute-resolution” practice area to test the chronology skill of the open-source (Apache 2.0 licensed) Claude-for-Legal toolkit across various Anthropic models and effort levels. This article turns to LAB itself, attempting to reproduce several aspects of the initial results Harvey published last week. As it developed, this article’s study diverged from the initial LAB results — and that divergence, rather than a clean reproduction, became its most useful result. Follow-up articles will examine how task-specific and practice-area-specific results were influenced by the structure and methodology of LAB.
Methods: Building a Proxy to Reproduce the Initial LAB Results
Harvey’s release anticipated that lawyers and researchers in the legal-AI community would join the conversation around LAB by running benchmark analyses, testing tasks, and reporting results under comparable conditions. The initial LAB results provide a natural baseline for that kind of comparison. This article attempts to participate in the LAB conversation by performing a study designed to reproduce aspects of the initial LAB results. Where it reproduces those results, it provides a baseline for more detailed analysis of them; where it diverges, it provides a means to investigate how the structure and methodology of LAB affect the scoring of evaluated models.
Ten AI models are evaluated by the initial LAB results on a set of legal tasks grouped into three practice areas: regulated and emerging company work; corporate transactions and funds; and privacy, tax, and private client work. Four models are scored within those practice areas: Opus 4.7, Sonnet 4.6, GPT-5.5, and Gemini 3.5 Flash.
[Image copied from Initial Results on Legal Agent Benchmark]
Because the initial LAB results do not disclose which specific tasks populate each practice-area grouping, this article cannot run the same evaluation exactly. Instead, it constructs a proxy by running six evaluated AI models against six representative LAB tasks. The evaluated models included Opus 4.7, Sonnet 4.6, GPT-5.5, and Gemini 3.5 Flash, along with Gemini 3.1 Pro and the newly released Opus 4.8. Two tasks were chosen to represent each of the three groupings for a total of six tasks.
For corporate transactions and funds, the representative tasks are “draft diligence summary memo” and “review commercial contracts diligence.” For regulated and emerging company work, the representative tasks are “triage vendor ai contracts for compliance with eu ai liability directive” and “analyze eu ai act high.” For privacy, tax, and private client work, the representative tasks are “analyze cpra compliance gaps against current privacy program” and “analyze counterparty markup of data processing agreement.”
The evaluated models’ generated outputs were initially judged by five models. Sonnet 4.6 and GPT-5.4 did not provide reliable scoring,3 leaving three reportable judges: GPT-5.5, Opus 4.7, and Opus 4.8. As discussed below, those scores support comparison across tasks, practice-area groupings, and judge models.
Results: How the Proxy Analysis Scored AI Models
When scores from all three judges (Opus 4.7, Opus 4.8, and GPT-5.5) are averaged across all six tasks, Opus 4.8 and Opus 4.7 achieve the best per-criterion coverage of approximately 95%. In practical terms, the Opus models rarely missed more than a couple of rubric criteria on any task.
Sonnet 4.6 and GPT-5.5 formed a second tier, trailing the Opus models by roughly four to seven percent. Both met approximately nine of every ten rubric criteria, though this varied by task. Sonnet 4.6, for example, achieved the highest average “draft diligence summary memo” score (75/76) while placing second to last on the “triage vendor ai contracts for compliance with eu ai liability directive” task.
Gemini 3.5 Flash formed a third tier at roughly 80% per-criterion coverage, with especially task-dependent performance.4 It tied for the second-highest average “analyze counterparty markup of data processing agreement” score (57/59) but otherwise ranked fourth or fifth on the evaluated tasks.
Gemini 3.1 Pro ranked last on every task, at approximately 54% average per-criterion coverage and 24.6% on its weakest task. Based on this analysis, Gemini 3.1 Pro does not appear competitive with the other evaluated models on the selected LAB tasks.
These per-criterion results differ from Harvey’s reported baseline for LAB in ways the next section examines. Task-specific and practice-area-specific results will follow in subsequent articles; in the meantime, model outputs and judged rubric scores are progressively being uploaded to GitHub.
Discussion: Asking Why the Proxy’s Results Diverged from the Initial LAB Results
Despite attempting to reproduce aspects of the initial LAB results, the proxy produced different ranking and score patterns. That makes it less useful as a direct reproduction of the reported LAB baseline, but more useful as a way to inspect how LAB results may be shaped by task selection, scoring methodology, judge model selection, and aggregation under the all-pass standard.
The most visible example is the model ranking. Excluding models not reported by practice area, the initial LAB results and the proxy ranked the evaluated models in the same order for only one practice-area grouping: corporate transactions and funds, where both ranked them (1) Opus 4.7, (2) Sonnet 4.6, (3) Gemini 3.5 Flash, and (4) GPT-5.5.
The other two groupings diverged. The initial LAB results show different models outperforming one another across the three practice areas, whereas the proxy’s rankings were more consistent across them. Individual tasks still showed model-specific variation, but that variation tended to narrow with aggregation, even with only two tasks per grouping.
The most immediate explanation for the proxy’s deviation is that models were ranked by all-pass scoring in the initial LAB results, whereas the proxy ranked them by per-criterion coverage. These metrics answer different questions, so part of the divergence may be definitional rather than substantive. The two scores should correlate, since higher per-criterion coverage generally produces higher all-pass scores. However, all-pass scoring compresses per-criterion scoring into binary pass/fail results, which can obscure analytically important distinctions. The Opus models illustrate this point by achieving roughly 95% per-criterion coverage while failing to achieve all-pass on most tasks.
Beyond different scoring methodologies, the proxy’s divergence from the initial LAB results may be explained by task selection, judge model selection, and run variance:
Task selection may matter substantially. The initial LAB results report practice-area scores in 2.5-percentage-point increments, suggesting each grouping may include 40 tasks. If so, those groupings are necessarily more selective than the tasks in practice-area folders. The corporate-ma folder, for example, includes 136 tasks; even if the corporate transactions and funds grouping drew only from that folder, fewer than one in three of its tasks would be included. Tasks selected to represent the corporate transactions practice area may not have been part of the grouping reflected in the initial LAB results, and the same is true for the tasks chosen to represent the other two groupings.
Judge model selection can materially affect results. Across all six evaluated models, the three reportable judges marked criteria as passing at different rates: Opus 4.8 at 86.3%, Opus 4.7 at 85.4%, and GPT-5.5 at 81.5%. These are judge pass rates and should not be confused with the per-criterion coverage scores reported above. All three judges agreed on 92.4% of criteria, and the two Opus judges agreed on 98.7%. That leaves 7.6% of binary pass/fail judgments on which at least one judge disagreed, most often GPT-5.5. Furthermore, abandoned runs of Sonnet 4.6 and GPT-5.4 suggest that unreliable judge models can have outsized effects. These partially completed runs produced lower interim pass rates of roughly 67.5% for Sonnet 4.6 and roughly 61.3% for GPT-5.4, which would have weighted scores toward failure if included.
Run variance may matter at the margins. This study ran each model on each task once. Frontier models are not deterministic, so repeated runs may produce different outputs and different scores. Inter-run variance could contribute to task-level variation, but it is unlikely to explain all of the recurring divergence from the initial LAB results across multiple tasks and practice-area groupings.
No single one of these variables explains the divergence. Together, they show that assessment of evaluated models is highly dependent on methodological choices: which tasks are selected, how they are grouped, how they are scored, and which model does the judging. Because these choices appear to materially affect the assessment of evaluated models, they warrant further investigation.
Conclusion: A Benchmark, Not a Leaderboard
Harvey invited lawyers to help test LAB. This study set out to accept that invitation by reproducing aspects of the initial LAB results through approximation of as much of Harvey's methodology as could be derived from those results. The initial LAB results were not reproduced. Instead, the proxy’s divergence shows why LAB should be studied as a benchmark, not merely consulted as a leaderboard. Because task selection, and judge model selection can change which model ranks highest and whether an output passes at all, the methodology behind a LAB score deserves as much scrutiny as the score itself. Follow-up articles will undertake that investigation in task-specific and practice-area-specific context.
This analysis is for educational and informational purposes only. It is not legal, financial, or other professional advice and does not create an attorney-client or any other professional relationship. The analysis is based solely on publicly available information. Any views expressed are solely those of the author in the author’s individual capacity.
The Legal Agent Benchmark was accessed on May 27, 2026 at https://github.com/harveyai/harvey-labs (commit 01983e9).
The author is not affiliated with Harvey or any legal AI entity, has not communicated with Harvey, and has not used the Harvey platform. This article is derived from the author’s personal study of the LAB in support of their own efforts to build, evaluate, and improve AI pipelines for legal analysis.
Sonnet 4.6 was mechanically unreliable, looping more than five turns on many of its scoring assessments. GPT-5.4 was substantively unreliable, focusing on technicalities over substance (e.g., whether the generated output repeated rubric language word-for-word). Both issues would have artificially inflated per-criterion failure rates for the evaluated models.
At the time of writing, Opus 4.8 scoring of Gemini 3.5 Flash outputs remained in progress for the “analyze cpra compliance gaps against current privacy program” and “triage vendor ai contracts for compliance with eu ai liability directive” tasks. For Gemini 3.5 Flash, the results of those two tasks were calculated as an average of the GPT-5.5 and Opus 4.7 judge scores. All other reported scoring was calculated as an average of the GPT-5.5, Opus 4.7, and Opus 4.8 judge scores.




