Cross-model benchmark
Cross-Model Legal Benchmark
HAQQ's own measured benchmark scoring legal AI models out of 50 across legal task categories.
HAQQ measured 19 models across 11 legal task categories, each scored out of 50. HAQQ's own engine (Justinian) leads most categories, including 49/50 on the top-scoring tasks, but it's not a clean sweep: Spellbook edges it out on contract drafting (46 vs. 44), LexisNexis leads legal research (46 vs. 43), and ChatGPT scores highest on plain-language explanation (45 vs. 42). Practical takeaway: no single model dominates every task, so choice should follow the workflow, not a leaderboard rank.
I dati
| Model | Overall /50 | Generale | Redazione di contratti | Ricerca giuridica | Spiegazione delle leggi | Contratto di lavoro | Redazione di memorandum | Contratto di licenza | Patti parasociali | Contratto di consulenza | Contratto commerciale | Redazione di NDA |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HAQQ (Justinian) | 45.8 | 49 | 44 | 43 | 42 | 48 | 44 | 46 | 47 | 45 | 47 | 49 |
| Claude Fable 5 | 42.7 | 45 | 45 | 41 | 43 | 43 | 41 | 42 | 42 | 42 | 41 | 45 |
| Claude Opus 4.7 | 40.8 | 43 | 43 | 39 | 39 | 39 | 41 | 43 | 40 | 39 | 42 | 41 |
| Mike OS | 39.1 | 42 | 39 | 35 | 33 | 41 | 39 | 38 | 41 | 40 | 40 | 42 |
| Harvey | 37.9 | 38 | 40 | 32 | 27 | 42 | 34 | 39 | 44 | 38 | 43 | 40 |
| DeepSeek v4 Pro | 36.8 | 40 | 36 | 38 | 38 | 35 | 37 | 36 | 36 | 36 | 37 | 36 |
| CoCounsel | 36.6 | 37 | 38 | 40 | 31 | 37 | 42 | 35 | 37 | 34 | 35 | 37 |
| Legora | 35.9 | 33 | 42 | 27 | 26 | 40 | 30 | 40 | 39 | 41 | 38 | 39 |
| ChatGPT 5.5 | 35.3 | 39 | 34 | 34 | 45 | 33 | 35 | 32 | 32 | 35 | 34 | 35 |
| Claude + legal plugins | 33.9 | 35 | 35 | 33 | 35 | 34 | 33 | 34 | 34 | 33 | 33 | 34 |
| Gemini 3.1 Pro | 32.9 | 36 | 32 | 36 | 41 | 30 | 32 | 31 | 30 | 30 | 31 | 33 |
| Spellbook | 32.9 | 27 | 46 | 18 | 20 | 38 | 20 | 41 | 35 | 37 | 36 | 44 |
| LexisNexis +AI | 32.0 | 36 | 29 | 46 | 28 | 30 | 38 | 28 | 31 | 27 | 29 | 30 |
| Grok 4.3 | 30.7 | 33 | 31 | 26 | 36 | 28 | 29 | 29 | 28 | 31 | 35 | 32 |
| Perplexity Sonar | 27.2 | 29 | 22 | 43 | 34 | 24 | 28 | 23 | 23 | 25 | 24 | 24 |
| Clio Duo | 25.6 | 26 | 27 | 24 | 23 | 28 | 23 | 25 | 24 | 29 | 26 | 27 |
| Meta Llama 4 | 23.8 | 24 | 23 | 23 | 29 | 23 | 24 | 22 | 22 | 24 | 23 | 25 |
| Mistral 3 | 22.4 | 22 | 25 | 20 | 25 | 21 | 22 | 24 | 20 | 22 | 22 | 23 |
| Qwen 3 Plus | 18.7 | 19 | 18 | 21 | 22 | 17 | 18 | 18 | 17 | 19 | 18 | 19 |
Dati misurati - pre-elaborati dal dataset di utilizzo reale dell'Indice Legal AI di HAQQ (oltre 134.000 punti dati in 30 paesi) o dal benchmark legale cross-modello /50 proprietario di HAQQ.
Ricerche correlate
- Multi-Model Strategy in LegalHow many AI models firms run, and the criteria they weigh when choosing.
- Risk & Hallucination RatesEstimated error rates across legal tasks, highest for citation and jurisdiction work.
- Competitive Landscape MapLegal AI vendor categories, vendor counts and growth rates across the market.