CRUSETRA
Appendix B Security and data handling ← Back to the findings

What runs on your machine, and what reaches the network.

A matte aluminium padlock; its shackle, closed, is the same deep green as the site's accents.

This page lists each place the tool touches, on your machine and on the network: the one network call it makes, the three places it writes, the packages it ships, and what an audit cannot promise. When a claim comes from the source code, this page gives the file and line where you can check it.

One kind of network call, and a guard at each of the four places in the code that make it

When the tool runs a measurement, it makes exactly one kind of network call, to the local model runtime, which is Ollama. By default it reaches that runtime on this machine's loopback address, the address a machine uses to talk to itself. The source has four places that send a document to that runtime. Each of them first passes a guard, which rejects any host that is not this machine and writes the refusal in plain words.

A test enforces the guard: it reads each source file, finds each of those four places, and fails the suite if a fifth place appears without the guard. The only way around the guard is an explicit flag on the command line, where an auditor can see it.

Check the guard

Run the test suite. The test that reads the sources fails if one of the four places has lost its guard.

run it yourself
npm testfails if one of the four places loses its guard

Where it livessrc/tiers.ts:556 · src/frontiere.test.ts:33

The models are pinned and cached, and their weights take 1.3 GB

Inference runs locally, on four sets of model weights. The sets are pinned by revision and by SHA‑256 per file, and that hash is a fingerprint of the file's content. The weights are cached on disk, and the tool checks them before it writes anything. Two points concern a bank in particular.

  • On the first run, the tool fetches the weights once from huggingface.co, 1.3 GB in total, and it announces that size before it starts. If your network policy forbids that download, the weights can be exported on one machine and imported on yours. The OCR engine's two language files (5.2 MB, pinned by SHA‑256) come down the same way, on npm run tessdata -- --prime only; a reading never fetches them.
  • We found and closed one last call. Even with each weight cached, the model library used to ask huggingface.co for one tokenizer configuration file, the file that tells a model how to split text into pieces, each time the extractor models loaded. The library asks for that file at revision main whatever revision is pinned, and it looks in its cache before the network. The loader now places a copy of the pinned file under that cache key, so no call remains, whether the network is reachable or blocked. The offline mode is tested end to end: the suite runs a measurement with the offline flag set and requires it to finish without a single outbound request. If a model is missing, the tool stops before it attempts anything, with a refusal that gives its reason.
Check the weights

The pinned revisions and the per-file hashes live in the source. A test checks that the tool does not run on altered weights.

run it yourself
npm testdoes not run on absent or altered weights

Where it livessrc/poids.ts:68 · src/poids.test.ts:129

What the tool needs to run

The tool needs four things.

  • Node 24 or newer.
  • 1.3 GB of disk for the pinned model weights, fetched once or imported offline.
  • A local Ollama runtime for the generative readers, the language models that read your documents, reached on this machine's loopback address by default.
  • Enough free memory. The bench runs on modest machines, but below 1 GB of free memory it marks its own durations as non‑transportable, meaning not comparable with another machine's, and it does not publish figures it cannot stand behind.
Check the memory floor

The minimum Node version is written in package.json. The memory floor is written into each timing record that the bench holds back.

run it yourself
node --versionprints your Node version, to compare with the minimum above

Where it livespackage.json:60 · src/tiers.ts:295

The files it reads, and the three places it writes

The tool reads only the files you name by flag. It never scans a disk.

It writes to three places. The first is its own records, under its repository root, in a folder that git ignores, so nothing measured on your data can travel into a public repository. The second is one report deposited beside the file you pointed it at, named after that file, with its sealed record (hashed, then frozen: its content hash is checked before a figure is shown). The third is one local marker, the date of the tool's first use, in ~/.crusetra/premiere-utilisation.json (an earlier date under the former ~/.cascade still counts): one labeled date, which the thirty-day clause reads and which is never transmitted. Only one record is versioned, the egress observation, the record of the open connections the tool sampled while it ran. In that record your file paths are replaced by a placeholder before it is written, and a journal header records no machine name and no user name.

What the tool depends on

The tool has two production dependencies, the model runtime library and tesseract.js (the OCR engine, Apache-2.0), and two development dependencies. The lockfile lists 90 packages, each with its content hash. Of those, 64 install on a given machine, and the rest are per‑platform variants. A generator classifies their licenses, so no one sorts them by hand: 63 permissive, 1 with obligations, 0 blocking, 0 undetermined, re‑checked on each test run. A CycloneDX software bill of materials ships with the tool.

The repository itself has no install‑time script, a script that runs when a package is installed, beyond wiring its own git hooks. Two transitive production packages do carry install scripts. On macOS and Windows they touch nothing beyond the npm registry. On one platform, linux/x64, one of them fetches optional GPU providers from a second registry at install time. npm ci --ignore-scripts disables both scripts, and running on CPU, as documented, loses nothing.

Check the counts

The license table and the test count are both generated. Each has a checker that fails the suite when the document drifts from the sources.

run it yourself
node src/licences.ts --checkfails when the license table driftsnode src/readme.ts --checkfails when the test count drifts
Your reviewers can read the code

The license is PolyForm Noncommercial, which makes the code source‑available. Anyone may read, study and fork the code with no time limit. An evaluation grant lets any organization run it on its own data, internally, for thirty days. One limit remains: it is not an OSI‑approved license, and some organizations exclude such licenses by policy.

What this page cannot promise

The tool watches itself run. It samples open connections during a measurement and publishes what it saw. That record states its own limit. Sampling can catch a connection that was open. It cannot prove that none was. The guarantee on this page is the guarded code path and the test that checks it. The observation does not carry that guarantee.

A content hash cannot promise permanence either. A public record that carries its content hash proves what was measured on that day's cases and nothing beyond them. recertify measures again on fresh records, flags drift in the input, says for each field whether the result holds or moved, and re‑computes the report's content hash. It runs on the validity period you declare, every ninety days by default.

The figures on this page are checked the same way as the rest of this site. The sentence 867 tests across 100 files is generated from the sources and re‑checked on each test run, so a stale figure cannot pass unnoticed.

Measure the drift

Rerun on fresh records against the baseline and its content hash. The answer says which fields hold and which moved.

run it yourself
npm run recertify -- --cases=fresh.csv --baseline=your-file-measured.jsoncompares each fresh case with the baseline, answers per field, and computes the report's content hash again