Harvey Unveils Initial Results of Legal AI Benchmark LAB
Harvey has released the first results from its Legal Agent Benchmark (LAB), an open-source framework designed to evaluate AI agents on complex, long-horizon legal tasks. The initial findings, published May 26, 2026, underscore significant limitations in current-generation AI models. Despite rapid advancements, frontier models completed less than 10% of LAB tasks end-to-end under a strict all-or-nothing evaluation standard. LAB, launched earlier this month, evaluates AI agents across over 1,200 tasks spanning 24 legal practice areas. Each task mirrors real-world law firm workflows, requiring AI models to produce review-ready legal work products graded against 75,000 expert-created rubric criteria. Harveys “all-pass” scoring system demands perfection—every rubric criterion must be satisfied for a task to pass. Key Findings: Frontier AI Falls Short Among the evaluated models, Claude Opus 4.7 led with a 7.1% success rate, followed by Sonnet 4.6 at 5.4%, Opus 4.6 at 4.2%, GPT-5.5 at 2.1%, and Gemini 3.5 Flash at just 0.8%. While these figures suggest progress, they also highlight how far legal AI lags behind human capabilities. “Legal work is far from saturated,” the report notes, especially given the high stakes and precision required in domains like corporate law, IP, and regulatory compliance. The findings also revealed uneven competence across practice areas. Models