Cross-model benchmark
Cross-Model Legal Benchmark
HAQQ's own measured benchmark scoring legal AI models out of 50 across legal task categories.
HAQQ measured 19 models across 11 legal task categories, each scored out of 50. HAQQ's own engine (Justinian) leads most categories, including 49/50 on the top-scoring tasks, but it's not a clean sweep: Spellbook edges it out on contract drafting (46 vs. 44), LexisNexis leads legal research (46 vs. 43), and ChatGPT scores highest on plain-language explanation (45 vs. 42). Practical takeaway: no single model dominates every task, so choice should follow the workflow, not a leaderboard rank.
البيانات
| Model | Overall /50 | عام | صياغة العقود | البحث القانوني | شرح القانون | عقد العمل | صياغة المذكرات | اتفاقية الترخيص | اتفاقية المساهمين | اتفاقية الاستشارات | الاتفاقية التجارية | صياغة اتفاقية السرية |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HAQQ (Justinian) | 45.8 | 49 | 44 | 43 | 42 | 48 | 44 | 46 | 47 | 45 | 47 | 49 |
| Claude Fable 5 | 42.7 | 45 | 45 | 41 | 43 | 43 | 41 | 42 | 42 | 42 | 41 | 45 |
| Claude Opus 4.7 | 40.8 | 43 | 43 | 39 | 39 | 39 | 41 | 43 | 40 | 39 | 42 | 41 |
| Mike OS | 39.1 | 42 | 39 | 35 | 33 | 41 | 39 | 38 | 41 | 40 | 40 | 42 |
| Harvey | 37.9 | 38 | 40 | 32 | 27 | 42 | 34 | 39 | 44 | 38 | 43 | 40 |
| DeepSeek v4 Pro | 36.8 | 40 | 36 | 38 | 38 | 35 | 37 | 36 | 36 | 36 | 37 | 36 |
| CoCounsel | 36.6 | 37 | 38 | 40 | 31 | 37 | 42 | 35 | 37 | 34 | 35 | 37 |
| Legora | 35.9 | 33 | 42 | 27 | 26 | 40 | 30 | 40 | 39 | 41 | 38 | 39 |
| ChatGPT 5.5 | 35.3 | 39 | 34 | 34 | 45 | 33 | 35 | 32 | 32 | 35 | 34 | 35 |
| Claude + legal plugins | 33.9 | 35 | 35 | 33 | 35 | 34 | 33 | 34 | 34 | 33 | 33 | 34 |
| Gemini 3.1 Pro | 32.9 | 36 | 32 | 36 | 41 | 30 | 32 | 31 | 30 | 30 | 31 | 33 |
| Spellbook | 32.9 | 27 | 46 | 18 | 20 | 38 | 20 | 41 | 35 | 37 | 36 | 44 |
| LexisNexis +AI | 32.0 | 36 | 29 | 46 | 28 | 30 | 38 | 28 | 31 | 27 | 29 | 30 |
| Grok 4.3 | 30.7 | 33 | 31 | 26 | 36 | 28 | 29 | 29 | 28 | 31 | 35 | 32 |
| Perplexity Sonar | 27.2 | 29 | 22 | 43 | 34 | 24 | 28 | 23 | 23 | 25 | 24 | 24 |
| Clio Duo | 25.6 | 26 | 27 | 24 | 23 | 28 | 23 | 25 | 24 | 29 | 26 | 27 |
| Meta Llama 4 | 23.8 | 24 | 23 | 23 | 29 | 23 | 24 | 22 | 22 | 24 | 23 | 25 |
| Mistral 3 | 22.4 | 22 | 25 | 20 | 25 | 21 | 22 | 24 | 20 | 22 | 22 | 23 |
| Qwen 3 Plus | 18.7 | 19 | 18 | 21 | 22 | 17 | 18 | 18 | 17 | 19 | 18 | 19 |
بيانات مقاسة - تمت معالجتها مسبقًا من مجموعة بيانات استخدام مؤشر الذكاء الاصطناعي القانوني الحقيقية لـ HAQQ (أكثر من 134,000 نقطة بيانات عبر 30 دولة) أو معيار HAQQ القانوني الخاص بـ /50 عبر النماذج.
أبحاث ذات صلة
- Multi-Model Strategy in LegalHow many AI models firms run, and the criteria they weigh when choosing.
- Risk & Hallucination RatesEstimated error rates across legal tasks, highest for citation and jurisdiction work.
- Competitive Landscape MapLegal AI vendor categories, vendor counts and growth rates across the market.