CRUSETRA

See which engine each field of your documents actually needs.

A model tier is the size of model a field is sent to, from a plain text pattern up to the largest.
On our identity-record test set, three of the five fields are read by a text pattern alone, at no cost.

git clone https://github.com/ArslaneSempai-ui/crusetra-routing cd crusetra-routing node src/premiere-reponse.mjs Prints the receipts result from the signed CORD record, then our KYC corpus. Under one second, before npm install.
The measured relief, state 01: 94.4% per field, 76.7% per file published routing: 94.4% accuracy per fieldper-file rate: 76.7%, 92 of 120 fileshuman tier: 85% assumed, not sampled
finding 0117.7 points apart

The same published routing reads 94.4% accuracy averaged per field and 76.7% as the per-file rate, 92 of 120 files, 17.7 points apart.

94.4%mean accuracy over the five fields
76.7%files with all five fields correct (92 of 120)
The measured relief, state 02: File-aimed routing: 3 files gained, none lost, of 120 files file-aimed routing: cheaper only on an assumed pricepublished routing: the pick it replacesa tier both routings share
finding 02on an assumed price

Aiming at the file gains 3 files and loses none of the 120 files: too few to separate the two rates. The cost falls 3.5× only if the large model is billed at an assumed price per call.

$191Published routing, per 100,000 documents, with the large model at an assumed $1.60 per 1,000 calls.
$54File-aimed routing, same volume and assumption. Priced at machine time, it is the dearer of the two.
The measured relief, state 03: Abstention: 85 wrong values removed, 12 correct values withheld an emptied cell: a blank held for re-readingabstention: 85 wrong values removed, 12 correct withheld
finding 03after abstention

On 30 documents chosen for being hard, abstaining removes 85 values that were wrong and withholds 12 values that were right. Of the values still returned, 62.3% are right, against 30% before.

30%Accuracy when each value is returned, right or wrong, on the 30 documents of the hard corpus.
62.3%Accuracy of the values still returned after abstention; 97 of 150 values go to review.
The measured relief, state 04: Identical counts across two passes, and unstable durations withheld two runs, identical countsdurations vary 16 to 60%, so withheld
finding 04withheld

Two passes produce identical counts. Durations vary, so they’re withheld.

identicalToken counts match exactly across both runs.
16–60%Durations moved between 16% and 60% from one run to the next, so they are withheld; a cost built on a duration carries that spread.
The measured relief, state 05: All 16,807 routings computed, none sampled all 16,807 routings tested, none sampledhuman tier: 85% assumed, not sampled
finding 05the full span

The solver computes all 16,807 routings and prints the winner; the record carries the count. None is sampled.

16,807All 16,807 routings computed, end to end.
120 filesThe 120 records held out for scoring them, frozen with a content hash.

The extraction cost audit

Find the cheapest extractor for each field of your documents, measured on your own pages.

  1. You label a sample of your own documents: for each field, the value it should read, or a dash when the document has no such line.
  2. You run the audit on your machine. Your current extractor, the ones you want to compare, and our local models read the same pages. The local models read them as text: the text your current vendor already returns, or an OCR you run.
  3. You receive a sealed record and a signed report: for each field, the cheapest source that stays within the margin you declare of the best, and the saving at your volume. Sealed means the record carries a content hash, so an edit made after sealing shows. Signed means the report page carries a signature your audit team checks against our public key with node src/verifier-rapport.mjs, like the sample below.
First page of the sample report: 100 real receipts, 8 sources on 3 fields, all routed to gemini-flash, saving $95,710 a year at 1,000,000 documents.
sample report · 100 real receipts · sealed 2026-09-29 · record ac7d0adbe4907caf · signed page beside the PDF

Start with receipts

The public run grades two vendors and our local tiers on 100 real receipts. Grade one vendor's outputs against the labels yourself, with nothing downloaded:

$ git clone https://github.com/ArslaneSempai-ui/crusetra-routing $ cd crusetra-routing && npm ci --ignore-scripts $ npm run grade -- --cases=examples/cord-receipts/cord-labels-grouped.csv --name=google-expense --values=examples/cord-receipts/cord-google-values.json --price-per-thousand-documents=100 --out=/tmp/google-expense.json

On those receipts, gemini-flash reads 96.8% of totals right and google-expense 93.7%. Our two encoder tiers read 2.1% and 62.1%; the local tier that competes, gen-4b, reads 81.1%, and it needs Ollama. The local tiers read a text that our macOS OCR produced from the images, so their rates include that OCR's errors.

Snapshot

$490up to 1,000 pages

One document type, measured once and sealed, back within 48 hours.

  • Up to 10 fields and 4 extractors
  • The routing for each field, and the saving at your volume
  • What the sample is too small to decide

Audit

$1,900up to 10,000 pages

Up to 3 document types, sealed, back within 72 hours.

  • Up to 30 fields and 10 extractors
  • A sample sized to separate sources a few points apart, when they are
  • The signed report, its PDF, and the sealed record

Quarterly audit

$4,900a year

The audit measured again 4 times a year, as vendors change their models and prices.

  • 4 sealed reports a year
  • Each one says what moved since the last
  • Stop at any time

The measurement runs on your machine and offline. One sealed record comes to us: counts, a right, wrong or blank verdict per case and field, your file's name and its hash, and the prices and volume you declared. No document, no value read from one. Testing a cloud extractor sends pages to that vendor under your own account, as it does today; the local tiers keep each page on your machine. Vendor fees for the pages you run through cloud extractors are billed to you by those vendors, on your own keys.

How the free test goes

  1. Label 100 pages of one document type in a CSV: an id, the text, then one column per field, each named with its type, as in id,text,total:amount,date:date,vendor:free-text. Without a type the comparison is exact text, so "$1,234.50" and "1234.50" count as different; the tool refuses such a column unless you pass --exact. Write a dash where a document has no such line, and leave the cell empty when you do not know the value. For the text, take what your vendor already returns, npm run text-from-exports -- --cases=your.csv --vendor=<textract|documentai|azure> --exports=<folder>, or read your images offline: once, npm run tessdata -- --prime (5.2 MB of language files), then npm run text-from-images -- --cases=your.csv --images=<folder> --ocr=tesseract --lang=eng. Either one writes your-with-text.csv: use that file in steps 2 and 3. Node 24 or newer. Up to 3 extractors.
  2. Grade each extractor you use or want to compare. We call no vendor: you run each one on your pages, then grade its outputs here. From a JSON of values, { "<id>": { "<field>": "<value>" } }: npm run grade -- --cases=your-with-text.csv --name=<vendor> --values=<its outputs> --price-per-thousand-pages=<your price> --out=<vendor>.json. From raw Textract, Document AI or Azure exports: --vendor=<textract|documentai|azure> --exports=<folder> --mapping=mapping.json in place of --values. The file it writes holds verdicts, no value.
  3. Measure, with the margin you accept: npm run measure:yours -- --cases=your-with-text.csv --sorties=<vendor>.json,<vendor>.json --current=<the one you run today> --margin=2 --pages-per-year=<your volume>, one --sorties per file or a comma list. Add --no-encoders to compare vendors only, with no download; without it, our local models download once.
  4. Email the record, the file ending in -measured.json beside your CSV, to contact@crusetra.com. It holds counts, a right, wrong or blank verdict per case and field, your file's name and its hash, and the prices you declared: no document, no value. It is the one file you may attach; a price under a vendor contract can be replaced with a list price before you send it.
  5. You get a one-page PDF back by email within 48 hours, from that record. The record carries a content hash; this report is not signed, and the free test stays an internal evaluation under the thirty-day grant. The Snapshot and the Audit come back signed, and include the right to act on the recommendation in your own operations. On our 100-receipt sample, two vendors 3.2 points apart on totals could not be told apart (7 disagreements, p = 0.45): 100 pages separate vendors far apart, and the paid tiers size the sample for close ones.

Download the sample report Verify the signed report See the labels, the OCR text and the vendor outputs Open the sealed record

Try the Routing instrument on our public test set.

Crusetra · RoutingOpen the live instrument See every field, tier, accuracy and cost, live from our public test set. Set the budget and watch the tool choose.
See pricing

Measured on 1,000 held-out records for the rules, small and large tiers, and 120 for the generative tiers, with a ring on the published routing of each field. The records are a corpus we wrote, about 166 characters each; the generative tiers' prompt was tuned on the held-out half, so their figures are optimistic by an amount not yet measured. Human accuracy is assumed at 85 % until you measure your own reviewers.