In public
Accuracy
How EXISTS, SAYS and HOLDS score against a benchmark of real disputed citations, our worst number first. Updated by the nightly benchmark run; a share we mark INSUFFICIENT_EVIDENCE is reported as such, never folded into a pass.
not enough answered to score
no method answered on enough of the benchmark to score it — insufficient evidence on EXISTS 61% of 197, SAYS 62% of 123
Citations in the Wild v0.2 · 13 documents · 197 claims · generated 2026-09-11
case registry · open search 105 · search by name 66 · no answer 26
By level
| level | n | balanced accuracy | INSUFFICIENT_EVIDENCE share | 95% interval |
|---|---|---|---|---|
| EXISTS | 197 | — | 61% | — |
| SAYS | 123 | — | 62% | — |
| HOLDS | 0 | — | — | — |
INSUFFICIENT_EVIDENCE share this month: 100%
Month by month
| month | receipts | claims | citations | not checked | INSUFFICIENT_EVIDENCE (of citations) | EXISTS fail (of claims) | SAYS fail (of claims) |
|---|---|---|---|---|---|---|---|
| 2026-09 | 2 | 2 | 2 | 0 | 100% | 0.0% | 0.0% |
How to read this
EXISTS and SAYS are near-deterministic: a source resolves or it does not; the attributed words are in it at a character offset or they are not. Their errors are fetch errors and drift, not judgement.
HOLDS scores 64–77 % on hard subsets. It disagrees with experts often, returns INSUFFICIENT_EVIDENCE often and on purpose, is off by default, and never enters the sealed record.
External reference numbers for the same class of task are in the low-to-mid 70s.
EXISTS asks the case registry two ways: an exact citation lookup, metered by the day, and past the meter the registry's open search, a looser instrument. How a run split between the two is printed under the benchmark line.
updated nightly by the benchmark run; a drift over 3 points changes this page automatically