[August 15, 2026] Harvey corrected the task documents identified in this article on July 30, 2026.
[July 27, 2026] While preparing for retrieval augmented generation (RAG) of Harvey’s new M&A Due Diligence tasks, a document ingestion pipeline returned a regulatory license that contained no license:
That is the entire tennessee-wholesale-distributor-license-tn-wd-729048.docx document of the pharma-pipeline-acquisition task. A separate ingestion run on a different task surfaced the same thing:
That is the entire sec-closure-letter-workplace-conduct-disclosure-investigation.docx document of the gaming-strategic-acquisition task. In total, 40 such artifacts were identified across nine of Harvey’s M&A Due Diligence tasks. Those familiar with vibe coding (or even just AI chatting) will recognize these as AI refusals, generated when a model is asked to do something that its programming prohibits.
On July 17, Harvey extended its Legal Agent Benchmark (LAB) to include new diligence tasks meant to represent the virtual data rooms (VDRs) common to M&A due diligence.1 The stated goal of these new tasks is to create “realistic, high-scale agent environments to evaluate agents’ ability to perform end-to-end legal work and support open model training and agent research.” While traditional LAB tasks contain about a dozen documents and perhaps 100,000 tokens, the new diligence tasks run to thousands of documents and as much as 80 million tokens.
Further, Harvey describes these larger diligence tasks as RL environments that will “support open model training and agent research,” including by providing “[e]nvironments to train agents for diligence” that “allow [them] to train diligence agents that can take an entire dataroom, client communication and deal context and product [sic] a diligence memo” by “scal[ing] RL + post-training to handle this scale and complexity.”
Large data sets benefit from techniques like retrieval augmented generation (RAG) because those sets exceed the context windows of frontier models. RAG can resolve such issues by providing AI agents with the ability to search, filter, identify, and collect the most pertinent context from a larger set. Harvey’s announcement prompted “excite[ment] to test this larger dataset by giving frontier models a queryable RAG agent”:
However, the AI refusal artifacts indicate that LAB’s new diligence tasks may not be ready for RAG and that AI-generated content may have been published without verification.
First, AI refusals themselves are less remarkable than their presence in the preeminent benchmark for the accuracy of legal AI analysis. The publication of AI refusals without verification is consistent with the kind of autonomous AI generation (i.e., generation without an attorney in the loop) that is irking courts and leading to fines and suspensions. Indeed, an attorney who files 40 AI refusal artifacts in a brief or diligence report is likely to face sanctions. Even more concerning, such artifacts appear throughout tasks designed to train agents for M&A diligence through reinforcement learning (RL), for which accuracy and comprehensiveness seem paramount. If attorneys are sanctioned for relying on LAB-trained agents, will Harvey take responsibility?
Second, LAB’s realism should be tested rather than assumed. Apart from AI artifacts (e.g., refusals), the diligence sets are unusually uniform and exceptionally easy to parse. They are limited to five top-level file types (.docx, .xlsx, .pptx, .eml, .txt), with no PDFs, no comments (other than .pptx speaker notes), no hyperlinks, no MIME attachments, no images, no macros, and no scanned or OCRed documents. Consistent with generation by an AI model, LAB documents appear to have similar structure, voice, tone, and syntax. Attorneys may wish to ask whether that resembles the legal data sets (e.g., data rooms) that they’re familiar with before trusting LAB to evaluate models or train agents.
This article is part of a larger series describing the author’s experience “validat[ing] and improv[ing] the benchmark” pursuant to Harvey’s invitation in the announcement of LAB.2 A first article attempted to reproduce Harvey’s initial LAB results and instead found them highly dependent on methodological choices. A second article sought to isolate the scoring variance associated with grading judge model selection, found that this variance rendered all-pass scoring unreliable, and advocated for per-criterion scoring. A third article identified hallucinated scoring criteria while running skill-matched Claude-for-Legal assessments on LAB tasks. RAG ingestion of LAB’s diligence tasks identified the AI refusals described here.
In view of repeated shortcomings, LAB’s standing as the preeminent legal AI benchmark may be worth reconsidering. Identifying these AI refusals required nothing more than a parser or a basic grep. A benchmark that skipped such nominal diligence should not own the conversation about legal AI, where diligence, accuracy, and verification matter.
Methods: The Case for Retrieval Augmented Generation (RAG)
Large data sets are common in several legal practice areas, including M&A diligence and litigation discovery. RAG is a well-established technique for mitigating context-window limits and context bloat by providing AI agents (e.g., frontier AI models in an agent harness) with the ability to query external indexes and retrieve the most pertinent chunks of context for a specific analysis. This is not a comprehensive review of RAG, which encompasses numerous approaches: VectorRAG, GraphRAG, hybrid search, re-ranking, iterative agentic retrieval, and structure-aware chunking. The recently published “Retrieval-Augmented Generation for Natural Language Processing: A Survey” provides a good overview of RAG:
Shangyu Wu et al., “Retrieval-Augmented Generation for Natural Language Processing: A Survey,” Artificial Intelligence Review (June 1, 2026).
For LAB’s diligence tasks, the attempted implementation was a structure-aware vector RAG pipeline. First, source documents (.docx, .xlsx, .pptx, and .eml, with .txt pass through) were parsed into candidate structural units using a custom ingestion pipeline. Qwen3.5-9B then adjudicated those units by determining whether they should remain separate, be merged, or be subdivided into retrieval chunks. The RAG pipeline was interrupted here due to Qwen3.5-9B’s identification of AI refusal artifacts while adjudicating chunk boundaries.
Had the RAG pipeline not been interrupted, the resulting chunks would have been embedded into vector indexes using Qwen3-Embedding-8B, then stored in a SQLite-backed Chroma database. Retrieval would have run through a query tool built on Qwen3-Embedding-8B and Qwen3-Reranker-8B, allowing frontier models (e.g., Opus, Fable, GPT-5.6, Grok 4.5) operating in an agent harness (e.g., Claude Code, Codex, Grok Build) to agentically query the RAG system for the most pertinent context.3
RAG was abandoned once the AI refusal artifacts indicated LAB’s diligence tasks were not ready for analysis. As detailed below, at least 40 refusal artifacts sit in plain document body text, readily identifiable by any parser (including a basic grep) and visible to anyone who opened the files or extracted their contents. Apparently, no one did.
Results: At Least 40 AI Refusal Artifacts
Once the tennessee-wholesale-distributor-license-tn-wd-729048.docx of the pharma-pipeline-acquisition task and the sec-closure-letter-workplace-conduct-disclosure-investigation.docx of the gaming-strategic-acquisition task were identified, the full diligence folder of the LAB repository was scanned. That scan identified the following AI refusals.
The aerospace-vertical-integration task includes at least 9 AI refusal artifacts:
prairie-aerostructures-holdings-inc-delaware-certificate-of-good-stand.docx
prairie-aerostructures-uk-ltd-companies-house-certificate-of-good-stan.docx
prairie-aerostructures-holdings-inc-kansas-foreign-corporation-status-.docx
prairie-aerostructures-uk-ltd-companies-house-current-filing-status-ce.docx
780-max-fuselage-traveler-msn-7110.docx
780-max-recovery-traveler-msn-7190.docx
as9100d-recertification-audit-report-april-2022-prairie-opco-wichita-t.docx
qp-0704-odar-02-authorized-inspection-delegation-and-stamp-control-rev.docx
wichita-building-4-plug-door-verification-record-780-qa-2406-122.docx
The cybersecurity-tuck-in task includes at least 1 AI refusal:
ipr2024-00313-exhibit-1001-u-s-patent-no-11-502-847.docx
The enterprise-software-diversification task includes at least 1 AI refusal:
scdf-fire-certificate-cecil-court-18th-floor-novavault-singapore-offic.docx
The gaming-strategic-acquisition task includes at least 8 AI refusal artifacts:
sec-closure-letter-workplace-conduct-disclosure-investigation.docx
confetti-mobile-ltd-texas-foreign-registration-and-franchise-status-ce.docx
meridian-ireland-ltd-certificate-of-good-standing.docx
meridian-studios-llc-delaware-certificate-of-good-standing.docx
palisade-interactive-inc-delaware-certificate-of-good-standing.docx
palisade-japan-kk-certificate-of-registered-matters.docx
accc-informal-merger-review-closure-letter-northwind-palisade.docx
cma-provisional-findings-issues-letter-original-transaction.docx
The grocery-horizontal-merger task includes at least 4 AI refusal artifacts:
cottonwood-companies-inc-california-combined-corporation-franchise-tax.docx
cottonwood-specialty-rx-llc-delaware-certificate-of-good-standing-july.docx
cottonwood-companies-inc-delaware-certificate-of-good-standing-july-20.docx
dea-registration-certificate-and-renewal-file-cottonwood-pharmacy-az-3.docx
The media-recap task includes at least 3 AI refusal artifacts:
2100-riverside-parkway-atlanta-production-center-annual-fire-operation.docx
501-north-state-wmdn-tv-chicago-studios-occupancy-and-fire-safety-cert.docx
5555-melrose-studios-culver-city-annual-fire-clearance-2024.docx
The pharma-pipeline-acquisition task includes at least 5 AI refusal artifacts:
tennessee-wholesale-distributor-license-tn-wd-729048.docx
bothell-fei-3004512890-fda-establishment-registration-2022-confirmatio.docx
kb-6091-ind-154288-may-proceed-letter.docx
kb-6814-ind-160455-may-proceed-letter.docx
companies-house-certificate-of-good-standing-kestrel-biotherapeutics-u.docx
The restaurant-pe-buyout task includes at least 7 AI refusal artifacts:
norrell-stanton-bank-statement-and-reconciliation-bsi-operating-accoun.docx
norrell-stanton-bank-statement-boardwalk-gift-card-settlement-account-.docx
norrell-stanton-bank-statement-boardwalk-global-marketing-coop-ad-fund-010.docx
norrell-stanton-bank-statement-bsc-boardwalknext-fee-collections-accou-024.docx
norrell-stanton-bank-statement-fwh-franchise-royalty-collections-accou-022.docx
norrell-stanton-bank-statement-fwh-ifpc-rebate-receipts-account-octobe.docx
thames-mercantile-bank-statement-boardwalk-uk-operating-account-octobe.docx
The sports-media-consolidation task includes at least 2 AI refusal artifacts:
kentucky-sales-and-use-tax-return-and-eft-confirmation-q1-2022.docx
bloodline-live-chicago-2023-09-12-local-event-permit-packet.docx
This analysis is for educational and informational purposes only. It is not legal, financial, or other professional advice and does not create an attorney-client or any other professional relationship. The analysis is based solely on publicly available information. Any views expressed are solely those of the author in the author’s individual capacity.
The Legal Agent Benchmark was accessed on July 23, 2026, at https://github.com/harveyai/harvey-labs (commit 845a088).
The author is not affiliated with Harvey or any legal AI entity, has not communicated with Harvey, and has not used the Harvey platform. This article is derived from the author’s personal study of the LAB in support of their own efforts to build, evaluate, and improve AI pipelines for legal analysis.
The parsing, embedding, and retrieval phases of this pipeline have been implemented previously on legal data sets; the author has not previously implemented structure-aware chunking on legal data sets.















































