The 100 real receipts of the CORD v2 test split (Clova AI, CC BY 4.0), with CORD's own labels; the split has been public since 2022, so the vendors' models may have seen it. Both vendor prices were declared by Crusetra for this sample. File cord-labels-grouped.csv, 100 cases, SHA-256 e8ff6b84e865bef90dbc477a39727a78…. Measured 2026-09-29 22:49 UTC with tool commit 7e82950, 8 sources on 3 fields. Every figure below is read from the sealed record ac7d0adbe4907caf (a content hash).
| field | send to | right, 95 % interval, cases | today | against google-expense, your current chain |
|---|---|---|---|---|
| total | gemini-flash | 96.8 % [91–99], 95 cases | 93.7 % | too close to call: 5 of 7 disagreements, p = 0.453 |
| subtotal | gemini-flash | 95.4 % [87–98], 65 cases | 89.2 % | too close to call: 6 of 8 disagreements, p = 0.289 |
| tax | gemini-flash | 97.5 % [87–100], 40 cases | 70 % | better: 12 of 13 disagreements, p = 0.003 |
| source | cost | total | subtotal | tax |
|---|---|---|---|---|
| gemini-flash | $4.29 per 1,000 documents | 96.8[91–99] | 95.4[87–98] | 97.5[87–100] |
| google-expense current | $100 per 1,000 documents | 93.7[87–97] | 89.2[79–95] | 70[55–82] |
| gen-4b | local, 804 ms a value | 81.1[72–88] | 78.5[67–87] | 27.5[16–43] |
| gen-0.6b | local, 315 ms a value | 77.9[69–85] | 40[29–52] | 27.5[16–43] |
| rules | local, under 1 ms a value | 66.3[56–75] | 56.9[45–68] | 67.5[52–80] |
| large | local, 21 ms a value | 62.1[52–71] | 32.3[22–44] | 10[4–23] |
| gen-8b | local, 1,162 ms a value | 61.1[51–70] | 50.8[39–63] | 35[22–50] |
| small | local, 14 ms a value | 2.1[1–7] | 0[0–6] | 0[0–9] |
At 1,000,000 pages a year (1 page a document), your current chain google-expense costs $100,000; the routing above costs $4,290. The saving, $95,710 a year, rests on these inputs: google-expense $100 per 1,000 documents, declared by Crusetra, for this sample; gemini-flash $4.29 per 1,000 documents, declared by Crusetra, for this sample; volume 1,000,000 pages a year, declared. Local sources: machine time at $1.20 an hour, an assumption; declare yours and the report follows. Accuracy and the verdicts are measured. The money is not: change a price and the figure moves with it.
These picks win on this sample, but the sample is too small to tell them apart; a larger sample could reverse them: gemini-flash against google-expense (your current chain) on total (95 cases, 7 disagreements, p = 0.453); subtotal (65 cases, 8 disagreements, p = 0.289).
Each source's answer is graded against the expected value, case by case, read as an amount written with grouped thousands (60.000 is sixty thousand). A case with no expected value is graded by nobody (total 5, subtotal 35, tax 60). Two sources are compared on the same cases: the cases where exactly one is right decide (McNemar, exact), and the worst case at 95 % is Newcombe's paired bound. A cheaper source is picked only if it cannot be more than 2 points worse than the best, the margin you declared. No rate is quoted under 20 cases. The local sources were asked a question made from each column name, such as “What is the total?”; your own wording can change their rates. The vendors' outputs are graded as they came.
It does not check a price against an invoice: every price is the one declared above. It does not see a value: the record holds right, wrong or blank per case, never what was read. It measures these 100 cases; documents of another kind, layout or language need their own sample. Local sources were timed on the machine that ran the audit; another machine takes another time.