Overfitting Dicta

Overfitting Dicta

Home
Notes
Twitter
Archive
About
Introducing Behavioral Analysis for Legal-AI
What session trajectories reveal about model behavior on twelve legal tasks
Sep 2 • Overfitting Dicta

July 2026

Investigating LAB, Part 4: RAG Impeded by the AI Refusals of Vibed Data Rooms
If Harvey plans to train agents in AI-generated RL data rooms, will those M&A diligence agents be able to identify AI refusal artifacts in their…
Jul 27 • Overfitting Dicta
Investigating LAB, part 3: Fabricated Rubric Criteria
LAB scoring rubrics include fabricated criteria possibly hallucinated during development
Jul 14 • Overfitting Dicta
Introducing a Twelve-Task Claude-for-Legal Benchmark
Defining a benchmark that tests twelve Claude-for-Legal skills from six practice areas on the Harvey LAB task that best exercises each skill
Jul 9 • Overfitting Dicta

June 2026

Investigating the LAB: Part 2, Judge Variation and the Case for Per-Criterion Scoring
How judge behavior affects the reliability and precision of legal AI assessment
Jun 9 • Overfitting Dicta
Investigating the Legal Agent Benchmark (LAB): Part 1
LAB should be studied as a benchmark, not merely consulted as a leaderboard
Jun 2 • Overfitting Dicta

May 2026

Benchmarking Claude for Legal
Testing Anthropic’s Claude for Legal toolkit against a litigation chronology benchmark
May 22 • Overfitting Dicta
© 2026 Overfitting Dicta · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture