CRUSETRA
Appendix A Method and reproducibility ← Back to the findings

How each figure is measured and then frozen.

A matte aluminium balance with two pans at exactly the same height; the right pan is the site's deep green.

This page shows the whole method: what is measured, on which slice of the corpus, how the routing is chosen, and how our public test set is frozen. That test set is the file of measured results we froze, and the code that reads it checks its content hash first. Each claim cites a path in the source that your reviewers can open. The steps also rerun on your own machine, so you can make the comparison yourself.

The figures come from a slice that the tuning did not see

The corpus is cut into three slices. One trains, one tunes (the dev slice), and one is held out from the tuning. The published measurement reads only the held‑out slice, and the code enforces that rule, so no discipline is needed. The dev slice exists because of a real mistake. A generative prompt, the text that instructs a model, was once tuned while it read held‑out scores. The leak was caught and repaired, and the split was hardened so it cannot recur unnoticed. A test fails the suite if the same document or the same wording appears in two slices.

One limit is stated by the repository itself: we wrote the corpus. The held‑out slice guards against grading cases the tuning had seen. It does not guard against cases that are unrealistic, which is why measuring your own cases matters.

Check the split

Run the suite. The split test goes through the corpus and fails if a document or a wording appears in two slices.

run it yourself
npm testincludes the split test

Where it livessrc/corpus.ts:164 · src/crusetra.test.ts:48

Each possible routing is tested, and the human tier is assumed

Each of the five fields can be routed to one of seven tiers. The optimizer tests all 16 807 possible routings, the complete set, by recursion. It uses no heuristic and no sampling. The count on this page is recomputed from the sources at each check. It is not typed in by hand.

The seventh tier, human review, is an assumption. It is not a measurement. The tool prints exactly that on each pass, and keeps the assumed figure apart from the six measured ones.

You can replace the assumption with a measurement. measure:humans grades your own reviewers on your own records. It reports accuracy per field with a confidence interval, the range the true rate is likely to sit in, agreement between reviewers, and seconds per record. Pass the result file, once hashed, to optimise with --humans, and the optimizer reads your measurement in place of the assumption.

Check that all routings are tested

Run the optimizer. It reads our public test set, goes through the routings, and flags the human tier as an assumption each time.

run it yourself
npm run optimiseprints the optimal routing and flags the human tier as assumednpm run measure:humans -- --cases=your-file.csvmeasures your reviewers: accuracy per field, agreement between them, seconds per record

Where it livessrc/optimise.ts:379 · src/readme.ts:849

A content hash freezes our public test set

Our public test set carries a content hash. The hash is a SHA‑256 of the file in canonical form, shortened to sixteen hexadecimal characters. The code that loads the file stops if the content hash is missing or no longer matches. It fails and names the cause, so a measurement edited by hand fails loudly. The stamp in this site's masthead is that content hash.

You do not have to take our word for it, and you should not hash the file by hand. A plain checksum of the file will not match, because the raw bytes are not what gets hashed. The content hash is computed on the canonical form of the file. The tool recomputes it for you and prints what it finds.

The hashing command states two limits about itself. The content hash covers the whole file except the content hash keys themselves. It also proves nothing about what happened before the hash was taken. Only the measurement command produces a measurement. One repair is also published in the open. The git commit at which the reference file was measured was later rewritten out of the repository history in a purge. The repository records the commit that stands in for it and labels that link “inferred, and not recorded”.

Recompute the content hash

Point the hashing command at the public test set shipped with the repository. It recomputes the content hash and prints it. On an untouched file it changes nothing.

run it yourself
npm run sceller -- profiles-2026-08-20-coeur-rendu.json --checkalready hashed, and the content hash matches: dbf26abec438515e

Where it livessrc/empreinte.ts:58 · src/sceller.ts:13

Each rate carries its confidence interval and its number of cases

Each rate comes with a confidence interval at 95%, computed by the Wilson method. The repository admits, in writing, that the level is an inherited default and that it was not weighed as a choice. The interval function throws an error on impossible input. It does not quietly return a NaN. For two tiers graded on the same cases, the repository's documented choice is a paired McNemar test. It calls a comparison of Wilson intervals too conservative for that job.

Check the arithmetic

The interval code is thirty lines, and its guard is tested. Our public test set records which interval method was used.

run it yourself
npm testcovers the interval guard and the stated method

Where it livessrc/interval.ts:32 · src/inventory.ts:63

Not every measured figure was published

On the constrained‑output bench, where the output is held to a set format, each duration moved by 16 to 60 percent between two passes on a memory‑starved machine. The durations were withdrawn, and in the data they were renamed so the defect is carried in the field name. The field is called msMedianeNonTransportableFamineMemoire, and the name says, in French, that the median milliseconds are non-transportable because of memory starvation. You cannot cite the field without its qualifier.

The token counts from the same bench reproduce to the digit, and those are published. A figure is shown only once it reproduces. Until then it is named in the data and kept off the page.

Check the retired figures

Rerun the bench. When the machine has less free memory than the bench needs, it marks its own durations non-transportable and leaves them unpublished.

run it yourself
npm run contrainteregenerates the bench file and qualifies the durations in the field name

Where it livesrapports/2026-08-22-contrainte-de-sortie.md:17 · contrainte.json:34

Run it twice and compare

The heading is an instruction. The full measurement reruns with one command. It requires an explicit flag, because the first run downloads 1.3 GB of pinned weights. The measurement itself took 32 minutes on the published pass. It then freezes a fresh result file you can set beside ours. Your own files are scored by the same grading code and the same intervals. Two frozen passes compare case by case, because a rate that rose may simply have lost cases.

A measurement also has a validity period. recertify reruns your measurement on fresh records against the frozen baseline. It checks the input for drift, that is, whether the fresh records differ from the baseline's. It returns one answer per field, holds or moved, and writes the report with a fresh content hash. You set the validity period, and the default is every ninety days.

Run it yourself

The four commands below are the whole loop: remeasure, measure your cases, compare, recertify.

run it yourself
MESURE_VOULUE=1 npm run measurereruns the full measurement and freezes a fresh test set with its content hashnpm run measure:yours -- --cases=your-file.csvscores your cases with the same grading code and intervals, and nothing leaves your machinenpm run diffcompares two frozen passes case by casenpm run recertify -- --cases=fresh.csv --baseline=your-file-measured.jsonmeasures fresh records against the frozen baseline, gives one answer per field, and hashes the new report

Where it livessrc/measure.ts:946 · src/your-cases.ts:1901