toolproof / index / doc-extract-bench|a masthead over nine indexes it did not buildmethodology 1.0 · CC-BY-4.0 · published by Compound Labs · read 13 Sept 2026 17:05 UTC
Nine indexes, one method9 of 9 answering, 1 behindAsk the nine
Sections

Nine indexes, one method

9 of 9 answering, 1 behind

9 indexes9 answering0 unreadable7 mint a badge2 deliberately not16 readings on 8 subjects751,973 listings read from source8.8% do not load

doc-extract-bench markdoc-extract-bench

Document-extraction APIs, scored on pinned public datasets.

On file, read on the capture date: obra/superpowers · anthropics/skills · n8n-io/n8n · anomalyco/opencode

the figure1,210documents scored
vendors compared3
pinned datasets5
extraction failures18
as the index dates itmeasured 2026-07-10This index publishes no staleness flag, so none is inferred here.
How it is testedthe index’s own sentence, as the read API returns it
doc-extract-bench markdoc-extract-benchmeasuredmeasured 2026-07-10

Every vendor is run against the same pinned subsets and EVERY RAW RESPONSE IS COMMITTED, so the published accuracy can be recomputed offline by anyone with no API keys and no account. It is the only one of the nine whose numbers a stranger can fully re-derive.

read from raw.githubusercontent.com/kyisaiah47/doc-extract-bench/main/results/latest.jsonthe index itself, github.com/kyisaiah47/doc-extract-bench

Measured 2026-07-10 by doc-extract-bench. Read into this register 13 Sept 2026. Toolproof measured nothing here.

What it can be asked aboutnobody, and that is deliberate

doc-extract-bench scores commercial extraction vendors on pinned public datasets. A vendor does not control the run and would be embedding a comparison against its competitors. The benchmark is published in full, every raw response is committed, and it is the only one of the nine a stranger can re-derive offline with no API keys at all.

Refusing to mint a badge that could only ever be used against its subject is the reason to trust the seven that are minted.

Why this index has no row on a repository8 subjects on file, all of them GitHub repositories

doc-extract-bench answers about nobody by design, so there is no per-subject reading to hold. Its figures are above, its method is above them, and the index is published in full at its own address.

Four kinds of answer, kept apart: a reading, an index that answered and holds nothing, an index that did not answer, and an index that was never the right one to ask. Collapsing the fourth into the second would be a statement about a subject that nobody made.