Every vendor is run against the same pinned subsets and EVERY RAW RESPONSE IS COMMITTED, so the published accuracy can be recomputed offline by anyone with no API keys and no account. It is the only one of the nine whose numbers a stranger can fully re-derive.
Measured 2026-07-10 by doc-extract-bench. Read into this register 13 Sept 2026. Toolproof measured nothing here.
doc-extract-bench scores commercial extraction vendors on pinned public datasets. A vendor does not control the run and would be embedding a comparison against its competitors. The benchmark is published in full, every raw response is committed, and it is the only one of the nine a stranger can re-derive offline with no API keys at all.
Refusing to mint a badge that could only ever be used against its subject is the reason to trust the seven that are minted.
doc-extract-bench answers about nobody by design, so there is no per-subject reading to hold. Its figures are above, its method is above them, and the index is published in full at its own address.
Four kinds of answer, kept apart: a reading, an index that answered and holds nothing, an index that did not answer, and an index that was never the right one to ask. Collapsing the fourth into the second would be a statement about a subject that nobody made.