Correctness Over Confidence: The 'I Don't Know' Test
In law, a beautifully wrong answer is worse than no answer. Fluency can hide fake law. The legal AI you can trust is the one that can say three words most systems can't: 'I don't know.' Judge tools on correctness, not confidence.
- Lesezeit: 11 min
Was dieses Kapitel behandelt
- Sounding Right vs Being Right
- The 1% That Matters
- Why AI Can't Be a Lawyer
- The Stockfish Standard
- The 'I Don't Know' Test
- Confirmation Bias
- How to Double-Check
Die Kurskapitel sind auf Englisch verfasst. Der Rest der Academy ist übersetzt.
TL;DR
AI is powerful because it is fluent, fast, and knows almost everything. That is also why it is dangerous. There is a wide gap between sounding right and being right, and in law the gap is where you lose. Pick the tool that values correctness over confidence - the one that can say "I don't know" instead of fabricating something that sounds smart.
1) Sounding right vs being right
AI is the best salesman you will ever meet. It is fluent, confident, and fast. But a good salesman makes you check whether the product is real. The gap between a fluent answer and a correct one is where legal work breaks. Fluency can hide fake law - an invented citation, a clause that reads clean and means nothing, a confidently wrong position.
2) The 1% that matters
In most fields, right 99% of the time is excellent. In law it is not enough, because the whole game is the 1% when the model is wrong and you cannot tell. You walk into a negotiation on a weak position without knowing it. The risk is not the average accuracy. It is the undetected error.
3) Why AI can't be a lawyer
"AI lawyer" is close to an oxymoron. Not because AI is weak - it can out-execute most people on speed, volume, and recall - but because a lawyer must be able to take responsibility and accountability. You cannot outsource judgment to a thing that cannot be held accountable. That is why a human stays in the loop, and why AI is a partner, not a source of law.
4) The Stockfish standard
Magnus Carlsen is the best chess player in history, and he will never beat Stockfish. Computers are simply better at the execution. People still play chess, still watch Carlsen, and still value the human. Law is the same - except the stakes are higher, so the human's job shifts to the part machines cannot own: judgment about what "winning" even means for this client, in this matter.
5) The "I don't know" test
There is a simple test for a legal AI you can trust. Can it say three words: "I don't know"? Most people can't, and almost no model can either - "I don't know" sits outside its parameters, so it guesses fluently instead. The first legal AI that will say "I don't know" when it doesn't know is the one you can trust, because it would rather stop than fabricate. Look for the tool that values correctness over confidence.
6) Confirmation bias
A warning most people miss: AI tends to complete you. Tell it "there's a problem with clause B" and it will find a problem with clause B, because agreeing finishes your thought. Separate your suspicion from its output, or you build a private echo chamber and get confidently misled. Ask neutrally. Let it disagree.
7) How to double-check
The check does not have to be a person - it can be another model. Run several AIs over the same document and compare. Direct the review where you are genuinely suspicious, not where you have already decided. Then let a human give the final judgment on what is right for this case.
Practitioner rule
Trust the tool that can say "I don't know." A beautifully wrong answer is worse than no answer, and in law, wrong is expensive.