CRUSETRA
Appendix A · Monitoring Method and reproducibility ← Back to the findings

What the Monitoring figures are made of.

A matte aluminium balance with two pans at exactly the same height; the right pan is the tool's deep lapis blue.

Before anything else: our public test set holds no real transaction. It is a file we froze, its content hash checked before any figure is shown. The measure that counts is yours, on alerts your analysts already closed, at your desk. Below: the cases, scenarios and scales behind each figure, and where one is not quoted.

Written cases and their benign counterparts

The measured half of our public test set rests on cases we wrote and mark as ours in the data: six suspicious typologies (structuring, rapid movement, dormant burst, round-tripping, cash-intensive, layering fan-out) and, deliberately, their six benign cases. Payroll moves money fast. A market season looks like a dormant account waking. A loan instalment repeats the same amount to the last unit. Rents cluster like structuring. What tells each benign case apart from its typology is in the data, history and validity period, and a test checks that structure. The label alone never carries it.

A set without these benign cases would flatter each scenario, because any scenario looks good when the benign cases are easy. The cases are what the measurement is for, and the findings above say which scenario fires on which case, as measured.

Verify the cases

The test suite validates each written case again, its nature, its window and its structure.

run it yourself
npm test84 cases validated again, case invariants under test

Where it livessrc/cas-etiquetes.test.ts:11 · src/cas-etiquetes.test.ts:41 · src/cas.ts:47

Generated variants, kept apart by construction

The generated half comes from the written cases under a stated seed, the fixed number that starts the random draw. The variation obeys one rule above all others. It must not undo the nature it varies. Structuring deposits stay under the reporting threshold. Round-trip amounts and loan instalments move by one common factor, so their equal amounts stay equal. Round savings amounts stay round. Each variant passes the case checks before it is added.

The two halves are never blended. The set holds them in two blocks, each with its source in the data. When a scenario scores better on the generated half than on the written one, the page shows that gap. It does not average it away.

Check the separation

Open the public test set. Each block carries its source in the data.

run it yourself
python3 -c "import json; d=json.load(open('releve-public.json')); print(d['authored']['provenance'], '/', d['synthetic']['provenance'])"prints: authored / synthetic

Where it livessrc/synthetic.ts:31 · src/synthetic.ts:51 · src/measure.ts:5

Scales are stated in one file, never buried in the code

Two scenarios need a scale to turn a raw number into a score: the amount scale and the velocity scale. Both are assumed values with stated bounds and a stated source, written in one file with the reason for each. Neither is a constant buried in the code. Change a scale and the public test set must be measured again and signed again, and the content hash makes a silent retune impossible.

Check the stated values

The scales file carries the value, the source and the bounds of each assumption.

run it yourself
npm testthe assumption tests check source and bounds

Where it livessrc/assumptions.ts:29 · src/assumptions.ts:76 · src/assumptions.ts:97

A confidence interval on each rate, and no number where the count is small

Each rate comes with its Wilson confidence interval at 95 percent (the range the true rate is likely to sit in) and the count it was measured on. Below twenty observations a rate is noted, and no number is quoted for it. Recall is the share of true suspicious cases a scenario catches. On your own history the tool refuses recall outright with fewer than five confirmed suspicious cases. It prints the refusal and no figure.

The threshold, the score at which a scenario fires, sits on one grid shared by each scenario, 0.50 to 1.00 in steps of 0.01. A case counts as an alert when its score is at or above the threshold. The comparison is inclusive, so the strictest threshold still means something for a scenario that scores exactly 1.

Check the refusals

Both cut-offs are written in the code, and the test suite exercises both.

run it yourself
npm testinterval and refusal tests run with the rest

Where it livessrc/interval.ts:102 · src/optimise.ts:170 · src/rapport.ts:25 · src/measure.ts:71