CRUSETRA
Appendix C Questions ← Back to the findings

The eight questions a bank's reviewers ask, and what the source does about each.

A matte aluminium key lying across the frame; its bow is a fat green ring.

Eight objections, the ones a bank's reviewers actually raise. None of the answers below is a promise. Each one describes a mechanism in the source, and the page names its file. Where a mechanism has a limit, the limit is here too.

“What if our fields are not your five?”

The fields come from your file. In the CSV you pass to the measuring command, the first two columns are an identifier and the input. Each later column is a field to extract, named by its header. No other setting is needed. The question put to the models comes from the column name, so a field new to this repository is asked about without anyone writing code.

Your own rules enter as regexes, text patterns, and become a tier measured like any other, one of the methods the tool ranks. Only regexes are accepted. Other rule code stays with you, and the tool reports that it cannot see it. The answers of a filled‑in questionnaire become the assumptions of the run.

Try it on one file

You pass one CSV with one flag. A second flag adds your regexes as a tier at zero cost.

run it yourself
npm run measure:yours -- --cases=your-file.csv --rules=rules.jsonmeasures your fields and your rules as tiers

Where it livessrc/your-cases.ts:18 · src/your-cases.ts:437

“Which models run, and which call out?”

Two extractive encoder models run inside the process, pinned by revision, which locks them to one version, plus two embedding models for classification. Three generative models run through a local Ollama, pinned by digest, a checksum, and optional. On your own file, measure:yours measures the two extractive tiers by default. It adds your regexes as a tier at zero cost when you pass --rules, and the three generative tiers when you pass --llm. Up to six tiers are measured. The human tier is an assumption until you measure it, which measure:humans does on your own reviewers. The only network call on the measurement path goes to the generative host, and the tool checks that it is local immediately before each call. Pointed at a remote machine, it will not start unless you pass the override flag yourself.

An egress run samples the open connections during a pass, and its file is published. One whole pass on client cases showed 0 connections over 21 samples. The file records its own limit. Sampling can miss a connection between two samples, so the count is a floor. There is no capture at kernel level.

Watch it run

The egress command samples the open connections during a pass, then publishes the hosts it saw and whether any connection went out.

run it yourself
npm run egresswatches the pass, and the committed file shows 0 connections

Where it livessrc/tiers.ts:556 · egress.json:12

“Do we need a connection?”

You need one for the first measurement, and for the public benchmark if you run it. The first measurement downloads 1.3 GB of pinned model weights from huggingface.co, and it tells you before it starts; reading images with the OCR engine needs two language files, 5.2 MB, that npm run tessdata -- --prime fetches once. After that, nothing goes out. One call used to remain: at each load the model library asked for a tokenizer configuration at revision main. The local cache now answers it, because the loader places the pinned file under the key the library reads. The test suite makes no download. The public benchmark is the one command that fetches a dataset. Nothing of yours goes up.

On a closed network, the weights export to a USB drive on one machine and import on yours. Each file is checked against its SHA‑256 checksum and its pinned revision before anything is written, tokenizer configuration included. Under the offline flag the encoder tiers run air‑gapped, with no network at all. This was measured, and a test holds it. When a model is missing, the tool names it and stops before anything is attempted. The tool does not break mid‑run.

Cross the air gap

Export on a connected machine, import on the closed one. The offline flag keeps the run off the network.

run it yourself
npm run poids -- --export /media/usb/crusetra-weightspacks the pinned weights for transportnpm run poids -- --import /media/usb/crusetra-weightschecks hash and revision before writing

Where it livesREADME.md:109 · README.md:120

“Who sees our records?”

Nobody sees them. Measurements on your data live in a folder that git ignores, so no measurement can travel into a repository. The file the tool hands back holds counts, a right, wrong or blank verdict per case and field, your file's name and its hash, and the prices you declared: no document and no value, and a test holds it. The suite plants sentinel values, runs the tool, and fails if any of them appears in the output. That test runs the real extractors, so it needs the model weights: it runs on the continuous-integration runners, where the weights are primed, and on your machine once npm run poids -- --prime has fetched them; before that it stands aside by name.

The same holds when the tool grades your own chain, the way you extract these fields today. The outcomes file carries a right or wrong for each answer, never the values.

Check the test

The sentinel test is part of the suite. It plants values, runs the tool, and fails if one of them leaks into the output.

run it yourself
npm testfails if a client value reaches the returned file (with the model weights on the machine)

Where it livessrc/crusetra.test.ts:3557 · src/your-cases.ts:1918

“What if the answer is: change nothing?”

Then that is what the report prints. Your current chain enters the measurement record as a tier and is ranked with the others. When it wins, the output prints that it wins outright on this sample. Under 20 observations the tool makes no recommendation.

The cheapest recommendation is an option from the start, and the repository points that way. Its own README shows the cheapest tier often suffices. The validation dossier has a mandatory What gets worse section. On the measured corpus, regexes already carry 3 of the 5 fields at zero cost.

“How long, and what lands on our machines?”

The evaluation grant runs 30 days on your own data, inside your organization. You install Node 24 or newer on macOS or Linux, then about 400 MB of packages and 1.3 GB of weights on the first measurement. On Windows the suite runs in continuous integration minus two test files, named with their reasons, and Windows is claimed only as far as that run is green. The published pass took 32 minutes of measurement on the machine named in our public test set. That test set is a file we froze, and its content hash is checked. The generative tier is 8 GB more, and optional.

Before any install, one script with no dependencies reads our public test set from a fresh copy of the repository. It prints the conclusion in under a second.

Read the answer in one second

On a fresh copy, before npm install, you can already reproduce the first answer.

run it yourself
node src/premiere-reponse.mjsprints the conclusion from our public test set, with no install

Where it livessrc/premiere-reponse.mjs:10 · README.md:259

“What happens to the tool after the engagement?”

The license has three levels, and the third is paid. Anyone may read, study, fork and use the code noncommercially, with no time limit. An evaluating organization runs it for 30 days on its own data. Commercial use is a separately negotiated license, so the public license is not what a client buys. The repository states that this is not an OSI‑approved license (OSI is the Open Source Initiative). Some organizations bar such licenses from adoption without a contract, which has no effect on a negotiated engagement.

After you fork, the dependency surface is 63 permissive packages, 1 with obligations, 0 blocking. Permissive means the license asks only that you keep its notice. The one obligation applies only if the tool ships as a closed binary. It is not shipped that way.

“Why should we believe these figures?”

Because they are recomputed, locked by a content hash, signed, and retracted when wrong. The test suite first checks that the README, the record behind the landing page and the validation dossier still match the code. It then runs 867 tests across 100 files, counted from the sources. The security and license documents are generated, marked as such, and fail the suite when they drift. Each measurement record carries a content hash, and a hand‑edited one is refused. Reports are signed, the public key lives in the repository, and a script there checks any report against it without depending on us. An example report ships in the repository as rapport-exemple.html, so the check can be run before anything is bought. It was issued with that key on our own held-out corpus, the cases we keep apart for testing.

A retraction file lists each published conclusion that turned out to be wrong. The secret sweep goes through 633 commits and looks for 20 forms of secret. It reports 0 undeclared secrets reachable from HEAD, the current state of the repository.

Verify a report

Your audit team checks any signed report against the repository's public key. The check does not need us.

run it yourself
node src/verifier-rapport.mjs rapport.htmlproves origin and integrity, and says nothing about correctnessnpm testrecomputes each published figure against the sources

Where it livessrc/verifier-rapport.mjs:20 · README.md:236