This article investigates how different models score Harvey’s Legal Agent Benchmark (LAB) when used as judges.1 Part 1 of this investigation set out to reproduce aspects of Harvey’s initial LAB results and instead found that LAB assessments depend on methodological choices, including that judge model selection can materially affect results.2 This article continues that investigation by characterizing how six models behave as LAB judges: their overall pass rates, whether their scoring varies by evaluated model, and how often they agree with one another.
To isolate the effects of judging, this article holds constant the previous article’s tasks, evaluated models, and generated outputs. Those are six LAB tasks selected to represent the three practice areas defined by the initial LAB results, with each task run by six evaluated AI models to generate 36 outputs. Those 36 outputs were scored against the appropriate LAB rubric by six judge models: GPT-5.4, GPT-5.5, Haiku 4.5, Sonnet 4.6, Opus 4.7, and Opus 4.8. Scoring was performed at extra high (xhigh) effort by Opus and GPT judges and at high effort by Sonnet and Haiku judges (for which extra high effort is not an option). Because the same outputs, rubrics, and evaluated models were scored by each judge, differences in scoring isolate the act of judging itself.
Based on the following results, LAB scores do not represent model performance, they represent a judge’s perspective of model performance according to a rubric. Judge variation affects model scoring. Under all-pass scoring, this variation appears to overshadow about 40% of the all-pass assessments of evaluated models on LAB tasks. That renders all-pass scoring unreliable because it measures the judge nearly as much as the evaluated model. This article therefore abandons efforts to reproduce Harvey’s initial LAB results because, in light of observed judge variation, they cannot be reproduced. Instead, this article concludes that per-criterion scoring is the superior LAB evaluation methodology.
Results: Judge Behavior
The six task rubrics comprised 375 criteria (about 63 per task). Each judge made the same 2,250 score determinations by applying these 375 criteria to the six evaluated models. Because each judge scored the same generated outputs against the same criteria, differences in 2,250 determinations are attributable to judging.
Overall pass rates clustered in a narrow band
Across 2,250 score determinations, overall pass rates ranged from 80.8% for GPT-5.4 to 86.3% for Opus 4.8.
This distribution reveals permissiveness gradients across iterative generations of the same model, and for Anthropic by model tier. Within the GPT model family, overall pass rate increased 0.7% from 5.4 (80.8%) to 5.5 (81.5%). Within Opus, they increased 0.9% from 4.7 (85.4%) to 4.8 (86.3%). Meanwhile, overall pass rates increased 5% across Anthropic model tiers from the throughput-focused Haiku 4.5 (81.3%) to the flagship Opus 4.8 (86.3%). This may indicate that more advanced reasoning results in higher pass rates, but does not necessarily mean that more permissive judges are more accurate.
Per-criterion rankings were consistent, no significant biases observed
Judges did not demonstrate significant bias by evaluated model and per-criterion scores were stable.
None of the models demonstrated significant bias or self-judging preference. Generations of the Opus and GPT model families demonstrate consistent scoring with earlier generations judging 0-1% lower than newer generations. The only notable variation is Opus’s relatively higher GPT-5.5 pass rate, but this doesn’t appear to indicate concerning bias. Haiku is remarkably consistent with GPT models despite the previously described variation in individual score determinations. Consistent with broad cross-model alignment, Sonnet split the difference.
Agreement was high within families, lower across them
Generations of the Opus and GPT model families showed high scoring alignment. For example, Opus 4.7 and 4.8 agreed on 98.7% of scoring determinations, while GPT-5.4 and 5.5 agreed 96.5% of the time. Sonnet achieved the broadest cross-model alignment, averaging 96.05% consensus with Opus models, 94.25% consensus with GPT models, and 94% consensus with Haiku. Meanwhile, cross-model agreement between Opus, GPT, and Haiku falls to 92% (92.2-92.7% when Opus/GPT generations are averaged).
Although 92% alignment might appear good, the converse 7.4-7.8% disagreement is more consequential where overall pass rates are compressed into a relatively narrow 5.5% band (80.8-86.3%). In this context, alignments of less than 94.5% indicate that variance is not isolated to marginal pass determinations. Rather, there is cross-model discord in which higher pass rate judges fail criteria that stricter judges pass and vice versa.
Where judges split, the splits were structured
Across 2,250 score determinations, the judges universally agreed that an evaluated model passed a task criterion for 1,694 determinations (75.3%), universally agreed on failure for 260 determinations (11.6%), and disagreed in 296 split determinations (13.2%).
Split determinations generally followed the judges’ overall permissiveness and alignment patterns. They were structured, not random.
Opus judges formed a permissive bloc, in which one or both Opus judges voted pass in 85.1% of splits. GPT judges formed a strict bloc, in which one or both GPT judges voted fail in 68.9% of splits. Sonnet rarely dissented. Meanwhile, Haiku-4.5 was the most discordant judge, appearing as a lone dissenter in 63 splits.
Discussion: Judge Variation Makes Per-Criterion Scoring More Reliable Than All-pass Scoring
Judge behavior has implications for the most appropriate scoring methodology. For example, Opus, GPT, and Haiku disagreed with each other on an average of 4.72 criteria per evaluated model per task. Sonnet disagreed with GPT on 3.58 criteria per task and with Opus on 2.47. Even GPT-5.4 and GPT-5.5 disagreed on 2.19 criteria per task. That variation affects model scores. In this analysis, for example, GPT-5.5 was graded 11% higher by Opus 4.8 (91.5%) than by Haiku 4.5 (80.5%).3 Under per-criterion scoring, that variation remains visible because it can be traced to specific disagreements over specific criteria. Here, overall rankings remained stable across all six judges despite shifts of up to 11% in per-criterion scores.
All-pass scoring does the opposite. It magnifies judge variation by turning criterion-level disagreements into complete task-level failures. Had all-pass scoring been used for this analysis, model ranks would have been unstable and scores unreliable: some judges would have passed no models on any task, while others would have passed different models on different tasks.
This instability becomes more pronounced when near-passes are considered. Excluding Gemini 3.1 Pro, the analysis produced 180 judge–model–task evaluations under the all-pass rule: 5 evaluated models × 6 LAB tasks × 6 judges. Of those, 70 failed by fewer than 4.72 criteria, the observed average criterion-level disagreement among Opus, GPT, and Haiku. Thus, for competitive models, 38.9% of evaluations scored under the all-pass rule fell within the empirical margin of judge variation. This article uses per-criterion scoring because model-task evaluations should not pass or fail based on judge variation.
According to Harvey’s release of LAB, the case for all-pass scoring is grounded in how legal work is used.
The evaluation methodology for LAB puts it even more bluntly:
This article doesn’t dispute Harvey’s reasoning that incomplete legal analysis presents material risks or the LAB’s objective of measuring complete legal analysis. Rather, it presents empirical results showing that the all-pass score does not accomplish that objective. All-pass scoring works against that objective because it lets methodological variance obscure the assessment of the evaluated model. For example, in this analysis, the all-pass scores of different judges deviated on multiple vectors: not only in how many tasks were passed by each evaluated model, but which models passed which tasks. Ultimately, the all-pass score appears to measure judging methodology as much as it measures the evaluated model's performance.
Conclusion: Judge Variation Favors Per-Criterion Scoring
Per-criterion scoring is superior due to the variation inherent in judge selection and other methodological choices. All-pass scoring does not reliably measure the quality or completeness of legal analysis. It compresses each task into a single pass-or-fail verdict that magnifies and obscures methodological variation. This manifests as unstable and unreliable results that could not be reproduced. Meanwhile, per-criterion scoring makes methodological variance visible and interrogable with higher precision and more reliable results. All-pass scoring should be abandoned in favor of per-criterion scoring for LAB analysis.
This analysis is for educational and informational purposes only. It is not legal, financial, or other professional advice and does not create an attorney-client or any other professional relationship. The analysis is based solely on publicly available information. Any views expressed are solely those of the author in the author’s individual capacity.
The Legal Agent Benchmark was accessed on May 27, 2026 at https://github.com/harveyai/harvey-labs (commit 01983e9).
The author is not affiliated with Harvey or any legal AI entity, has not communicated with Harvey, and has not used the Harvey platform. This article is derived from the author’s personal study of the LAB in support of their own efforts to build, evaluate, and improve AI pipelines for legal analysis.
The “per-model pass-rates, by judge” chart displays this disagreement as 10% due to rounding. The actual scores are Haiku 4.5: 0.805333 and Opus 4.8: 0.914667.









