Skip to content

Cross-model benchmark

Cross-Model Legal Benchmark

HAQQ's own measured benchmark scoring legal AI models out of 50 across legal task categories.

MeasuredUpdated 2026-07-08

HAQQ measured 19 models across 11 legal task categories, each scored out of 50. HAQQ's own engine (Justinian) leads most categories, including 49/50 on the top-scoring tasks, but it's not a clean sweep: Spellbook edges it out on contract drafting (46 vs. 44), LexisNexis leads legal research (46 vs. 43), and ChatGPT scores highest on plain-language explanation (45 vs. 42). Practical takeaway: no single model dominates every task, so choice should follow the workflow, not a leaderboard rank.

The data

Cross-Model Legal Benchmark - HAQQ's own measured benchmark scoring legal AI models out of 50 across legal task categories.
ModelOverall /50LegalContract DraftingLegal ResearchLaw ExplanationEmployment AgreementMemo DraftingLicense AgreementShareholder AgreementConsultancy AgreementCommercial AgreementNDA Drafting
HAQQ (Justinian)45.84944434248444647454749
Claude Fable 542.74545414343414242424145
Claude Opus 4.740.84343393939414340394241
Mike OS39.14239353341393841404042
Harvey37.93840322742343944384340
DeepSeek v4 Pro36.84036383835373636363736
CoCounsel36.63738403137423537343537
Legora35.93342272640304039413839
ChatGPT 5.535.33934344533353232353435
Claude + legal plugins33.93535333534333434333334
Gemini 3.1 Pro32.93632364130323130303133
Spellbook32.92746182038204135373644
LexisNexis +AI32.03629462830382831272930
Grok 4.330.73331263628292928313532
Perplexity Sonar27.22922433424282323252424
Clio Duo25.62627242328232524292627
Meta Llama 423.82423232923242222242325
Mistral 322.42225202521222420222223
Qwen 3 Plus18.71918212217181817191819

Measured data - pre-processed from HAQQ's real Legal AI Index usage dataset (134,000+ data points across 30 countries) or HAQQ's own /50 cross-model legal benchmark.

Back to the Legal AI Index report

Related research

All research graphs