Skip to content
HAQQ

HAQQ Labs

The HAQQ Legal Benchmark

Two evaluations. A 50-point task benchmark over the work a lawyer hands over, and a cross-domain study that places a legal-specialised engine against the frontier models on their own ground. Justinian leads 6 of the 6 measured practice areas below.

How it is built, and what it does not cover

By practice area, out of 50

Each practice area is the mean of the measured task categories mapped to it. Areas the run does not cover are not shown rather than shown empty.

Overall coverage

  1. 1HAQQ (Justinian)49/50
  2. 2Anthropic: Claude Fable 545/50
  3. 3Anthropic: Claude Opus 4.843/50
  4. 4Mike OS42/50
  5. 5DeepSeek: DeepSeek V4 Pro40/50

Corporate & Commercial

  1. 1HAQQ (Justinian)47/50
  2. 2Anthropic: Claude Fable 543/50
  3. 3Anthropic: Claude Opus 4.841/50
  4. 4Harvey Tenet41/50
  5. 5Spellbook40/50

Employment & Mobility

  1. 1HAQQ (Justinian)48/50
  2. 2Anthropic: Claude Fable 543/50
  3. 3Harvey Tenet42/50
  4. 4Mike OS41/50
  5. 5Legora40/50

IP & Technology

  1. 1HAQQ (Justinian)46/50
  2. 2Anthropic: Claude Opus 4.843/50
  3. 3Anthropic: Claude Fable 542/50
  4. 4Spellbook41/50
  5. 5Legora40/50

Advisory & Research

  1. 1HAQQ (Justinian)46/50
  2. 2LexisNexis +AI42/50
  3. 3Anthropic: Claude Fable 541/50
  4. 4Thomson Reuters CoCounsel41/50
  5. 5Anthropic: Claude Opus 4.840/50

Private Client

  1. 1HAQQ (Justinian)46/50
  2. 2OpenAI: GPT-5.6 Sol Pro45/50
  3. 3Anthropic: Claude Fable 543/50
  4. 4Google: Gemini 3.1 Pro41/50
  5. 5Anthropic: Claude Opus 4.839/50

Cross-domain, out of 100

Twenty-five benchmarks across legal, tax, journalism, general capability, and safety and values, drawing on Stanford LegalBench, Harvey LAB, VLAIR, ALARB, BigLaw Bench, CUAD, LegalCiteBench and LawBench. The full table sorts by any column and filters by domain.

Cross-domain benchmark, 100-point scale, by domain and benchmark
BenchmarkHAQQ (Justinian)Anthropic: Claude Opus 4.8Thomson Reuters CoCounselGoogle: Gemini 3.1 ProZ.ai: GLM 5.2OpenAI: GPT-5.6 Sol ProKimi: K3Anthropic: Claude Sonnet 5Snowdon 1.0-LargeQwen: Qwen3.7 PlusDeepSeek: DeepSeek V4 Pro
Overall average83.279.578.578.077.276.575.775.773.673.071.9
Legal
Stanford LegalBench85.781.882.384.382.982.383.381.482.878.876.8
Info. Retrieval57.251.253.653.853.653.853.047.951.249.051.5
Reasoning78.275.273.277.873.168.474.873.170.866.570.1
Classification74.770.570.974.270.270.071.670.469.868.767.8
Doc. Processing & RAG83.278.979.778.276.578.679.074.771.274.173.7
Summarisation90.088.390.089.490.086.387.784.887.982.589.1
Contract Under.77.474.471.075.973.674.977.269.671.768.468.3
Human Queries89.785.489.286.487.388.661.484.387.186.786.4
Deep Research90.990.888.980.689.085.989.886.078.487.384.0
Harvey LAB86.886.985.755.584.676.183.780.956.370.683.1
Legal average81.478.378.475.678.176.576.175.372.773.375.1
Tax
Deep Research86.385.882.476.083.580.985.182.070.478.780.7
Tax Q&A88.986.987.984.585.786.688.388.788.285.182.8
Tax average87.686.385.180.284.683.886.785.479.381.981.8
Journalism
Deep Research84.682.380.983.079.066.384.574.176.277.278.6
Journalism average84.682.380.983.079.066.384.574.176.277.278.6
General
Factuality82.971.473.782.666.268.369.559.871.776.269.1
Long Context75.675.275.375.075.970.073.570.774.653.069.5
Multilingualism85.883.278.485.781.782.784.979.574.277.678.4
Instr. Following91.786.191.484.889.489.085.886.887.385.689.2
Writing80.779.380.378.578.179.178.278.779.977.980.7
Reasoning76.273.768.474.867.371.676.266.866.866.665.0
General Agent89.283.489.184.477.575.661.074.187.185.572.0
Coding66.857.439.950.056.050.866.857.440.943.945.6
Maths98.798.794.097.491.297.796.285.995.594.994.9
General average83.178.776.779.275.976.176.973.375.373.573.8
Safety / Values
Political Neutrality97.982.897.393.391.083.833.582.385.351.531.3
Robustness78.578.260.365.350.468.372.777.142.266.738.1
Adversarial Testing95.993.493.687.378.895.981.1
Safety / Values average90.883.678.364.568.771.450.2

An empty cell means the model was not run on that benchmark: it is not a zero, and it is left out of that model's averages. Sort and filter it on /justinian.

Reading these numbers

HAQQ builds this benchmark, runs it, and competes in it. That is why the methodology is published and the per-task results are visible rather than summarised: the check available to you is inspection, not trust. Rows where HAQQ is not first are rendered exactly like rows where it is.

Methodology and limits · The engine it measures