In public

Accuracy

How EXISTS, SAYS and HOLDS score against a benchmark of real disputed citations, our worst number first. Updated by the nightly benchmark run; a share we mark INSUFFICIENT_EVIDENCE is reported as such, never folded into a pass.

not enough answered to score

no method answered on enough of the benchmark to score it — insufficient evidence on EXISTS 61% of 197, SAYS 62% of 123

Citations in the Wild v0.2 · 13 documents · 197 claims · generated 2026-09-11

case registry · open search 105 · search by name 66 · no answer 26

By level

levelnbalanced accuracyINSUFFICIENT_EVIDENCE share95% interval
EXISTS19761%
SAYS12362%
HOLDS0

INSUFFICIENT_EVIDENCE share this month: 100%

Month by month

monthreceiptsclaimscitationsnot checkedINSUFFICIENT_EVIDENCE (of citations)EXISTS fail (of claims)SAYS fail (of claims)
2026-092220100%0.0%0.0%

How to read this

EXISTS and SAYS are near-deterministic: a source resolves or it does not; the attributed words are in it at a character offset or they are not. Their errors are fetch errors and drift, not judgement.

HOLDS scores 64–77 % on hard subsets. It disagrees with experts often, returns INSUFFICIENT_EVIDENCE often and on purpose, is off by default, and never enters the sealed record.

External reference numbers for the same class of task are in the low-to-mid 70s.

EXISTS asks the case registry two ways: an exact citation lookup, metered by the day, and past the meter the registry's open search, a looser instrument. How a run split between the two is printed under the benchmark line.

updated nightly by the benchmark run; a drift over 3 points changes this page automatically