Overfitting Dicta
Subscribe
Sign in
Home
Notes
Twitter
Archive
About
Introducing Behavioral Analysis for Legal-AI
What session trajectories reveal about model behavior on twelve legal tasks
Sep 2
•
Overfitting Dicta
2
3
July 2026
Investigating LAB, Part 4: RAG Impeded by the AI Refusals of Vibed Data Rooms
If Harvey plans to train agents in AI-generated RL data rooms, will those M&A diligence agents be able to identify AI refusal artifacts in their…
Jul 27
•
Overfitting Dicta
9
2
Investigating LAB, part 3: Fabricated Rubric Criteria
LAB scoring rubrics include fabricated criteria possibly hallucinated during development
Jul 14
•
Overfitting Dicta
3
2
Introducing a Twelve-Task Claude-for-Legal Benchmark
Defining a benchmark that tests twelve Claude-for-Legal skills from six practice areas on the Harvey LAB task that best exercises each skill
Jul 9
•
Overfitting Dicta
5
3
June 2026
Investigating the LAB: Part 2, Judge Variation and the Case for Per-Criterion Scoring
How judge behavior affects the reliability and precision of legal AI assessment
Jun 9
•
Overfitting Dicta
1
Investigating the Legal Agent Benchmark (LAB): Part 1
LAB should be studied as a benchmark, not merely consulted as a leaderboard
Jun 2
•
Overfitting Dicta
1
3
May 2026
Benchmarking Claude for Legal
Testing Anthropic’s Claude for Legal toolkit against a litigation chronology benchmark
May 22
•
Overfitting Dicta
2
2
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts