This article defines a twelve-skill, task-matched Claude for Legal benchmark of those skills that are best exercised by Legal Agent Benchmark (LAB) tasks.1
Claude for Legal is Anthropic’s open-source toolkit of legal skills that plug into a harness (Cowork or Claude Code) running a native Claude model to provide structured protocols that guide the model through specified legal analyses.2 Claude for Legal works by generating one or more profiles (company profile, practice profile, matter profile, etc.) that provide organizational, practice, and/or matter-specific context for legal analyses invoked by any particular skill. A previous Benchmarking Claude for Legal article tested the chronology skill of Claude for Legal’s litigation plugin with eleven model/effort combinations against a single build litigation case timeline task from Harvey’s LAB. That article found the chronology scores scaled with model capability and effort from 40% coverage of the task’s rubric with Haiku 4.5 Low to 86% coverage with Opus 4.7 Medium, then plateaued at higher efforts. No model/effort combination achieved 100% coverage and three rubric criteria (C-027, C-028, and C-045) were not captured by any of the eleven combinations. This first benchmark left open whether breadth of coverage is a good proxy for legal-analysis quality or whether the chronology results could generalize to other Claude for Legal skills.
The Legal Agent Benchmark is an open-source repository of synthetic legal tasks published by Harvey.3 Tasks are associated with various legal practice areas (e.g., litigation dispute resolution) and defined by a prompt, a scoring rubric, and documents comprising a factual record. Each task is designed to simulate a legal analysis performed by attorneys within the associated legal practice area. Harvey published the first iteration of LAB on May 6, 2026, published initial LAB results on May 26, 2026, then expanded LAB to include in-house contracting tasks on June 12, 2026. A previous Investigating the Legal Agent Benchmark (LAB): Part 1 article attempted to reproduce Harvey’s initial LAB results and instead found that they depended on undisclosed methodological choices, including task selection, grader model selection, and scoring methodology. A subsequent Investigating the LAB: Part 2, Judge Variation and the Case for Per-Criterion Scoring article determined that grader selection alone could shift an evaluated model’s per-criterion score by up to 11 percentage points, and that under all-pass grading roughly 40% of competitive evaluations fell within the resulting margin of judge disagreement. That article advocated for per-criterion scoring, which is adopted by this updated Claude-for-Legal benchmark.
A Twelve-Skill, Task-Matched Claude-for-Legal Benchmark
Claude for Legal skills are divided into twelve practice areas, nine of which are fields of law with substantive legal skills. When mapped to the closest corresponding LAB tasks, six of those nine yielded strong skill-to-task matches: AI governance, commercial, corporate, intellectual property, litigation, and privacy.4 The benchmark comprises the two strongest skill-task matches from each of these six practice areas.5,6
Each skill-task pairing reflects the Claude for Legal skill best exercised by a corresponding LAB task. The benchmark runs each skill on its paired task and scores the resulting work product against that task’s published rubric, criterion by criterion. It has been run against the four current-generation Claude models that can execute Claude for Legal — Sonnet 5, Opus 4.7, Opus 4.8, and Fable 5 — each at high, extra high, and max effort.
Conclusion
This benchmark is defined independently of the models it evaluates and of how their outputs are graded. Current-generation Anthropic models Sonnet 5, Opus 4.7, Opus 4.8, and Fable 5 have been evaluated, and results will be presented for each skill-task pairing. Recently released Grok 4.5, Meta Spark 1.1, and OpenAI's GPT-5.6 Sol, Terra, and Luna are being evaluated as comparative baselines; none is an Anthropic model, so none can run Claude for Legal skills.
This analysis is for educational and informational purposes only. It is not legal, financial, or other professional advice and does not create an attorney-client or any other professional relationship. The analysis is based solely on publicly available information. Any views expressed are solely those of the author in the author’s individual capacity.
The author is not affiliated with Anthropic, Harvey, or any other AI entity. This article is derived from the author’s personal study of Claude for Legal and LAB in support of their own efforts to build, evaluate, and improve AI pipelines for legal analysis.
The Claude-for-Legal repository was accessed on June 11, 2026 at https://github.com/anthropics/claude-for-legal (commit 248331e) under the Apache-2.0 license. Publication of this benchmark was delayed in order to cover current-generation Anthropic models Opus 4.8, Fable 5, Sonnet 5 models as released during development.
The Legal Agent Benchmark was accessed on June 13, 2026 at https://github.com/harveyai/harvey-labs (commit 47deaa8) under the MIT license.
The benchmark does not include employment, product, and regulatory practice areas because no LAB tasks cleanly exercised their skills. Nor does it include law student, legal builder hub, and legal clinic practice areas because they lack substantive legal skills.
"Regulation Gap Analysis" appears under both AI governance and privacy because each practice area includes a distinct "/reg-gap-analysis" skill performing similar analysis for separate practice areas.
“Claim Chart” is a litigation skill that is counted as an intellectual property skill due to its application to an intellectual property task. Both litigation and IP plugins were installed while running this skill.


