Our third citation benchmark, and the hardest. Starting from Harvey Labs' open Legal Agent Benchmark (an MIT-licensed task set spanning 24 practice areas, built by an outside team), we had two frontier AI models draft 502 complete legal deliverables, then had Kingsfield rule every citation those drafts produced.
We seeded 1,506 fabricated citations across the 502 drafts; Kingsfield caught every one, with none slipping through. On the citations we have independently ground-truthed so far, it produced zero false-accepts, meaning it never blessed a bad citation, the failure that gets a lawyer sanctioned. The full ground-truth census of the drafts' own naturally-generated citations is still running; we will report that accuracy number here, with its coverage, when it is complete.
The drafters are independent of the judge. Two frontier AI models, not built or tuned by us, wrote the 502 deliverables from Harvey Labs' task instructions. That matters: the citations Kingsfield rules on were produced by the same class of tool a firm would actually use, not authored by us to be easy or hard.
Ground truth is established independently, three ways. Whether each cited case exists is checked against a public case index (CourtListener), which is objective and does not depend on our corpus. Whether a citation supports the point it is cited for is judged by a separate oracle model that wrote none of the drafts, reading the full authority text. And quoted language is matched mechanically against the real opinion. A citation is only counted correct when it is real and used correctly.
The planted fabrications are the fail-open test. Into the real drafts we seeded 1,506 citations that do not exist. A judge that gets lazy at scale, or that rubber-stamps a long document, would let some through. Kingsfield caught all 1,506, which is the evidence that it rules on every citation rather than sampling.
Our Damien Charlotin benchmark and LegalCiteBench hand the judge citations one at a time. This one hands it whole documents, dozens of citations deep, across two dozen areas of law it did not train for. It is closer to what actually lands on your desk: a long AI-drafted brief where one buried fabrication is all it takes. Catching planted fakes proves the judge does not fall asleep at scale; never blessing a bad cite proves it does not hand you false confidence.
No. Harvey Labs' Legal Agent Benchmark is an open, MIT-licensed task set published by an outside team (github.com/harveyai/harvey-labs). We use it only as a neutral source of realistic legal tasks. Kingsfield is the judge that rules on the citations; the drafters and the ground-truth checks are independent of it.
A citation we deliberately inserted into a real draft that points to a case that does not exist. Because we know in advance it is fake, it is objective ground truth: the judge should reject it every time. Across the 502 drafts we planted 1,506 of them, and Kingsfield caught all 1,506.
Because we hold ourselves to ground truth. The drafts' own naturally-generated citations have not all been independently adjudicated yet, so labeling any one of them a "correct catch" would be getting ahead of the evidence. When that census is complete, we will add the case-level detail here.
Kingsfield tells you whether the citations hold up. It checks citations, not the merits of your argument, and a pass is not a guarantee. You make the final call on every result and keep your Rule 11 duty. See Legal & limitations.