Benchmark · Harvey-Labs

Zero false-accepts across 502 AI-drafted legal documents.

Our third citation benchmark, and the hardest. Starting from Harvey Labs' open Legal Agent Benchmark (an MIT-licensed task set spanning 24 practice areas, built by an outside team), we had two frontier AI models draft 502 complete legal deliverables, then had Kingsfield rule every citation those drafts produced.

How Kingsfield ruled

502AI-drafted documents
24 practice areas
5,421Citations in those drafts
2,040Unique authorities ruled
1,506 / 1,506Planted fabrications caught
100%Planted fabrications caught · 0 fail-open
0False-accepts on ground-truthed cites

We seeded 1,506 fabricated citations across the 502 drafts; Kingsfield caught every one, with none slipping through. On the citations we have independently ground-truthed so far, it produced zero false-accepts, meaning it never blessed a bad citation, the failure that gets a lawyer sanctioned. The full ground-truth census of the drafts' own naturally-generated citations is still running; we will report that accuracy number here, with its coverage, when it is complete.

Honest scope. The two numbers above are what we can defend today: a full-scale fail-open test (every planted fabrication caught) and a zero-false-accept result on the citations we have ground-truthed. We are not yet publishing a single end-to-end accuracy percentage for this set, because the census of the drafts' own naturally-generated citations is not finished, and we would rather show nothing than a number we cannot stand behind. That is the same rule we apply to our other two benchmarks.

How the benchmark was built

The drafters are independent of the judge. Two frontier AI models, not built or tuned by us, wrote the 502 deliverables from Harvey Labs' task instructions. That matters: the citations Kingsfield rules on were produced by the same class of tool a firm would actually use, not authored by us to be easy or hard.

Ground truth is established independently, three ways. Whether each cited case exists is checked against a public case index (CourtListener), which is objective and does not depend on our corpus. Whether a citation supports the point it is cited for is judged by a separate oracle model that wrote none of the drafts, reading the full authority text. And quoted language is matched mechanically against the real opinion. A citation is only counted correct when it is real and used correctly.

The planted fabrications are the fail-open test. Into the real drafts we seeded 1,506 citations that do not exist. A judge that gets lazy at scale, or that rubber-stamps a long document, would let some through. Kingsfield caught all 1,506, which is the evidence that it rules on every citation rather than sampling.

Why this is the hardest of our three tests

Our Damien Charlotin benchmark and LegalCiteBench hand the judge citations one at a time. This one hands it whole documents, dozens of citations deep, across two dozen areas of law it did not train for. It is closer to what actually lands on your desk: a long AI-drafted brief where one buried fabrication is all it takes. Catching planted fakes proves the judge does not fall asleep at scale; never blessing a bad cite proves it does not hand you false confidence.

Questions

Is Harvey Labs your product?

No. Harvey Labs' Legal Agent Benchmark is an open, MIT-licensed task set published by an outside team (github.com/harveyai/harvey-labs). We use it only as a neutral source of realistic legal tasks. Kingsfield is the judge that rules on the citations; the drafters and the ground-truth checks are independent of it.

What is a "planted fabrication"?

A citation we deliberately inserted into a real draft that points to a case that does not exist. Because we know in advance it is fake, it is objective ground truth: the judge should reject it every time. Across the 502 drafts we planted 1,506 of them, and Kingsfield caught all 1,506.

Why not show individual cases, like the other benchmarks?

Because we hold ourselves to ground truth. The drafts' own naturally-generated citations have not all been independently adjudicated yet, so labeling any one of them a "correct catch" would be getting ahead of the evidence. When that census is complete, we will add the case-level detail here.

Does a clean result mean my brief is safe to file?

Kingsfield tells you whether the citations hold up. It checks citations, not the merits of your argument, and a pass is not a guarantee. You make the final call on every result and keep your Rule 11 duty. See Legal & limitations.