HAQQ Labs
The HAQQ Legal Benchmark
Two evaluations. A 50-point task benchmark over the work a lawyer hands over, and a cross-domain study that places a legal-specialised engine against the frontier models on their own ground. Justinian leads 6 of the 6 measured practice areas below.
By practice area, out of 50
Each practice area is the mean of the measured task categories mapped to it. Areas the run does not cover are not shown rather than shown empty.
Overall coverage
- 1HAQQ (Justinian)49/50
- 2Anthropic: Claude Fable 545/50
- 3Anthropic: Claude Opus 4.843/50
- 4Mike OS42/50
- 5DeepSeek: DeepSeek V4 Pro40/50
Corporate & Commercial
- 1HAQQ (Justinian)47/50
- 2Anthropic: Claude Fable 543/50
- 3Anthropic: Claude Opus 4.841/50
- 4Harvey Tenet41/50
- 5Spellbook40/50
Employment & Mobility
- 1HAQQ (Justinian)48/50
- 2Anthropic: Claude Fable 543/50
- 3Harvey Tenet42/50
- 4Mike OS41/50
- 5Legora40/50
IP & Technology
- 1HAQQ (Justinian)46/50
- 2Anthropic: Claude Opus 4.843/50
- 3Anthropic: Claude Fable 542/50
- 4Spellbook41/50
- 5Legora40/50
Advisory & Research
- 1HAQQ (Justinian)46/50
- 2LexisNexis +AI42/50
- 3Anthropic: Claude Fable 541/50
- 4Thomson Reuters CoCounsel41/50
- 5Anthropic: Claude Opus 4.840/50
Private Client
- 1HAQQ (Justinian)46/50
- 2OpenAI: GPT-5.6 Sol Pro45/50
- 3Anthropic: Claude Fable 543/50
- 4Google: Gemini 3.1 Pro41/50
- 5Anthropic: Claude Opus 4.839/50
Cross-domain, out of 100
Twenty-five benchmarks across legal, tax, journalism, general capability, and safety and values, drawing on Stanford LegalBench, Harvey LAB, VLAIR, ALARB, BigLaw Bench, CUAD, LegalCiteBench and LawBench. The full table sorts by any column and filters by domain.
| Benchmark | HAQQ (Justinian) | Anthropic: Claude Opus 4.8 | Thomson Reuters CoCounsel | Google: Gemini 3.1 Pro | Z.ai: GLM 5.2 | OpenAI: GPT-5.6 Sol Pro | Kimi: K3 | Anthropic: Claude Sonnet 5 | Snowdon 1.0-Large | Qwen: Qwen3.7 Plus | DeepSeek: DeepSeek V4 Pro |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Overall average | 83.2 | 79.5 | 78.5 | 78.0 | 77.2 | 76.5 | 75.7 | 75.7 | 73.6 | 73.0 | 71.9 |
| Legal | |||||||||||
| Stanford LegalBench | 85.7 | 81.8 | 82.3 | 84.3 | 82.9 | 82.3 | 83.3 | 81.4 | 82.8 | 78.8 | 76.8 |
| Info. Retrieval | 57.2 | 51.2 | 53.6 | 53.8 | 53.6 | 53.8 | 53.0 | 47.9 | 51.2 | 49.0 | 51.5 |
| Reasoning | 78.2 | 75.2 | 73.2 | 77.8 | 73.1 | 68.4 | 74.8 | 73.1 | 70.8 | 66.5 | 70.1 |
| Classification | 74.7 | 70.5 | 70.9 | 74.2 | 70.2 | 70.0 | 71.6 | 70.4 | 69.8 | 68.7 | 67.8 |
| Doc. Processing & RAG | 83.2 | 78.9 | 79.7 | 78.2 | 76.5 | 78.6 | 79.0 | 74.7 | 71.2 | 74.1 | 73.7 |
| Summarisation | 90.0 | 88.3 | 90.0 | 89.4 | 90.0 | 86.3 | 87.7 | 84.8 | 87.9 | 82.5 | 89.1 |
| Contract Under. | 77.4 | 74.4 | 71.0 | 75.9 | 73.6 | 74.9 | 77.2 | 69.6 | 71.7 | 68.4 | 68.3 |
| Human Queries | 89.7 | 85.4 | 89.2 | 86.4 | 87.3 | 88.6 | 61.4 | 84.3 | 87.1 | 86.7 | 86.4 |
| Deep Research | 90.9 | 90.8 | 88.9 | 80.6 | 89.0 | 85.9 | 89.8 | 86.0 | 78.4 | 87.3 | 84.0 |
| Harvey LAB | 86.8 | 86.9 | 85.7 | 55.5 | 84.6 | 76.1 | 83.7 | 80.9 | 56.3 | 70.6 | 83.1 |
| Legal average | 81.4 | 78.3 | 78.4 | 75.6 | 78.1 | 76.5 | 76.1 | 75.3 | 72.7 | 73.3 | 75.1 |
| Tax | |||||||||||
| Deep Research | 86.3 | 85.8 | 82.4 | 76.0 | 83.5 | 80.9 | 85.1 | 82.0 | 70.4 | 78.7 | 80.7 |
| Tax Q&A | 88.9 | 86.9 | 87.9 | 84.5 | 85.7 | 86.6 | 88.3 | 88.7 | 88.2 | 85.1 | 82.8 |
| Tax average | 87.6 | 86.3 | 85.1 | 80.2 | 84.6 | 83.8 | 86.7 | 85.4 | 79.3 | 81.9 | 81.8 |
| Journalism | |||||||||||
| Deep Research | 84.6 | 82.3 | 80.9 | 83.0 | 79.0 | 66.3 | 84.5 | 74.1 | 76.2 | 77.2 | 78.6 |
| Journalism average | 84.6 | 82.3 | 80.9 | 83.0 | 79.0 | 66.3 | 84.5 | 74.1 | 76.2 | 77.2 | 78.6 |
| General | |||||||||||
| Factuality | 82.9 | 71.4 | 73.7 | 82.6 | 66.2 | 68.3 | 69.5 | 59.8 | 71.7 | 76.2 | 69.1 |
| Long Context | 75.6 | 75.2 | 75.3 | 75.0 | 75.9 | 70.0 | 73.5 | 70.7 | 74.6 | 53.0 | 69.5 |
| Multilingualism | 85.8 | 83.2 | 78.4 | 85.7 | 81.7 | 82.7 | 84.9 | 79.5 | 74.2 | 77.6 | 78.4 |
| Instr. Following | 91.7 | 86.1 | 91.4 | 84.8 | 89.4 | 89.0 | 85.8 | 86.8 | 87.3 | 85.6 | 89.2 |
| Writing | 80.7 | 79.3 | 80.3 | 78.5 | 78.1 | 79.1 | 78.2 | 78.7 | 79.9 | 77.9 | 80.7 |
| Reasoning | 76.2 | 73.7 | 68.4 | 74.8 | 67.3 | 71.6 | 76.2 | 66.8 | 66.8 | 66.6 | 65.0 |
| General Agent | 89.2 | 83.4 | 89.1 | 84.4 | 77.5 | 75.6 | 61.0 | 74.1 | 87.1 | 85.5 | 72.0 |
| Coding | 66.8 | 57.4 | 39.9 | 50.0 | 56.0 | 50.8 | 66.8 | 57.4 | 40.9 | 43.9 | 45.6 |
| Maths | 98.7 | 98.7 | 94.0 | 97.4 | 91.2 | 97.7 | 96.2 | 85.9 | 95.5 | 94.9 | 94.9 |
| General average | 83.1 | 78.7 | 76.7 | 79.2 | 75.9 | 76.1 | 76.9 | 73.3 | 75.3 | 73.5 | 73.8 |
| Safety / Values | |||||||||||
| Political Neutrality | 97.9 | 82.8 | 97.3 | 93.3 | 91.0 | 83.8 | 33.5 | 82.3 | 85.3 | 51.5 | 31.3 |
| Robustness | 78.5 | 78.2 | 60.3 | 65.3 | 50.4 | 68.3 | 72.7 | 77.1 | 42.2 | 66.7 | 38.1 |
| Adversarial Testing | 95.9 | — | 93.4 | — | 93.6 | — | 87.3 | — | 78.8 | 95.9 | 81.1 |
| Safety / Values average | 90.8 | — | 83.6 | — | 78.3 | — | 64.5 | — | 68.7 | 71.4 | 50.2 |
An empty cell means the model was not run on that benchmark: it is not a zero, and it is left out of that model's averages. Sort and filter it on /justinian.