CRUSETRA
Appendix A · Scoring Method and reproducibility ← Back to the findings

How the scoring grid is measured.

A matte aluminium balance with two pans at exactly the same height; the right pan is the tool's deep amethyst.

This page shows what the scoring numbers are made of. It names the customer files, the factors, the tables of stated values, and where a figure is not quoted. Our public test set is a file we froze, and its content hash, a fingerprint of its bytes, is checked before any figure is shown. No real person is in it, and no official risk list was copied into it. A test enforces both refusals, so neither rests on discipline.

The customer files we wrote, pair by pair

The test set has a written half and a generated half. The written one rests on customer files we wrote and state as ours. Six escalated typologies come with six retained files, written on purpose to match them. Escalated means an analyst sent the file up for review, and retained is its clean counterpart. A family company with each ownership layer named is paired with the shell that hides them. A neighborhood shop whose turnover sits on its declaration is paired with one whose cash takings run past it. A salaried account at a ratio of one, observed flows over declared, is paired with a firm at a ratio of five. A foreign student at stipend scale is paired with a newly onboarded customer, opened remotely, at several times his declaration. The two files of a pair differ in the fields, the history and the ratios. The tag alone never separates them. A test checks that all six pairs keep this structure.

No official country list was copied, and a test makes that refusal structural, part of the test set's own shape. Escalated and retained files share their countries of residence, so no country code can stand in for a list. The geography factor scores zero on this test set, on both sides. The Scoring page above this one reports that zero as a finding and does not hide it.

Check the customer files

The suite rechecks each written customer file, the six pairs and the shared-country rule.

run it yourself
npm testchecks the 84 customer files again, each file's rules and the no-list rule

Where it livessrc/dossiers-etiquetes.test.ts:65 · src/synthetic.ts:29 · src/synthetic.ts:50

Each table is stated, and stated as an assumption

A scoring factor needs tables that say how much weight each activity, each product and each channel carries. Each table in this tool sits in one file with its source, unit and bounds, and each carries the same one-word status, assumed. That status is the point of Scoring. The tables your vendor calls calibration are values someone stated. The only way to know what they are worth is to measure them against your own reviews that already reached a decision.

Check the assumptions

The tables file carries a value, a source and bounds for each assumption.

run it yourself
npm testthe assumption tests check source and bounds

Where it livessrc/assumptions.ts:81 · src/assumptions.ts:113

Generated variants are reported apart, whichever side they favor

The generated half comes from the written customer files under a stated seed, the number that fixes the random draw. The generator is built so that it cannot change a file's nature. Factors that depend on each other keep their observed-over-declared ratios. A file's cash share stays on the side of the limit it started on. Ownership layers keep their minimum count. The newly onboarded, remotely opened customer keeps that status. On this tool the generated half is harsher than the written one on the best factor, which means a lower figure there. It is reported apart whether it comes out worse or better. Separation is a rule. It is not a convenience. The page shows both numbers side by side.

Check the separation

Open the public test set. The two blocks carry their source in the data itself.

run it yourself
python3 -c "import json; d=json.load(open('releve-public.json')); print(d['authored']['provenance'][:24], '/', d['synthetic']['provenance'][:22])"prints the two source statements: written by hand, then seeded variants

Where it livessrc/synthetic.ts:29 · src/measure.ts:71

A confidence interval on each rate, scores kept within bounds, and no number when the count is small

Each rate travels with its 95 percent Wilson confidence interval, the range the true rate is likely to sit in, and the count it was measured on. Below twenty observations the rate is recorded but not shown. Recall on your own customer files, the share of your confirmed escalations it catches, is refused outright under five confirmed escalations. A factor's score is between 0 and 1, both ends included. A value outside that range is an error named with the factor and the customer file, never a silent rounding. The threshold grid, the set of score cut-offs, is shared and inclusive. Each cut-off includes its own value, so the strictest cell still means something for a factor that scores exactly one.

Check the refusals

The bounds and refusals live in code, and the test suite runs them.

run it yourself
npm testthe interval, score-bound and refusal tests run with the rest

Where it livessrc/interval.ts:102 · src/facteur.ts:55 · src/facteur.ts:56