CRUSETRA
Appendix A · Screening Method and reproducibility ← Back to the findings

How the Screening figures are measured.

A matte aluminium balance with two pans at exactly the same height; the right pan is the tool's deep ruby.

Here is what the Screening figures rest on: the pairs, the matchers (each one a way of comparing two names), the thresholds, and where a figure is not quoted. Each claim sits in the source at a path your reviewers can open, and each step reruns on your machine.

The written pairs, labeled by kind

Our public test set, a file we froze, has two halves, one measured and one synthetic. The measured half is 120 pairs written for this repository by an AI agent and labeled by kind (commit 474ffbd of 5 September 2026); no pair comes from a client or from a list, and the labels ship with the pairs so a debatable one can be argued on the record. Some are matches a screening system should catch: transliterations, changed word order, initials, typos, accents and marks, particles such as de or al, spellings by sound. Others are non-matching pairs it should ignore. The hard kind are near-matches, pairs that look alike but are not the same person, such as siblings, shared given names or near-identical spellings. The set counts each kind, so you can see how many of each it holds and which are missing.

Written pairs can only show so much, and the page with their figures repeats this next to each one. They prove what a matcher does on the variations they label, and they do not cover your alert history. The command that swaps our pairs for your history is one line, and it sends no data.

Check the test set

The test set is frozen (the tool's word is sealed). The repository's tool recomputes its content hash, a fingerprint of the file's bytes, and answers in one line.

run it yourself
npm run sceller -- releve-public.json --checksays: already sealed, and the seal matches

Where it livessrc/measure.ts:197 · src/measure.ts:123

The synthetic half, kept separate

The other half is generated by code: typos of four kinds, reversed word order, initials, stripped accents and marks, doubled letters, and a name transliterated more than one way. Each generated pair is labeled synthetic by kind and measured in its own tables, and they never enter a measured rate. For some matchers at some thresholds, the scores from which two names count as a match, the generated pairs score far better than the written ones, and a blend would flatter the tool just there.

The test set holds two blocks with two stated sources, and no code exists that could merge them. The synthetic generator is a separate module with a marked boundary, so when it is missing the measure still runs and reports absent for that half, and writes no empty half in its place.

Check the separation

Open the test set: each block carries its source in the data, in a field named provenance.

run it yourself
python3 -c "import json; d=json.load(open('releve-public.json')); print(d['authored']['provenance'], '/', d['synthetic']['provenance'])"prints: authored / synthetic

Where it livessrc/measure.ts:6 · src/measure.ts:127 · src/measure.ts:133

A confidence interval on each rate, and no number for a small sample

Each rate on the Screening pages comes with a 95 percent confidence interval by Wilson's method and the count of observations. The interval is the range the true rate is likely to sit in. Below twenty observations a rate is not quoted as a number, because the interval is then too wide to tell one result from another, and a printed number would claim more than the test set supports.

The same rule applies to recall on your own history. Recall is the share of true matches the tool catches. With fewer than five confirmed matches in your export, the tool says no confidence interval can be put on recall, prints no recall figure, and shows the synthetic tables on their own.

Check the refusals

Both cutoffs are set in the code, so this page only repeats them, and the test suite exercises them.

run it yourself
npm testruns the interval and refusal cases with the rest of the suite

Where it livessrc/interval.ts:102 · src/optimise.ts:183 · src/your-alerts.ts:162

One shared grid of thresholds, and equalling one counts as an alert

Each matcher returns a score between 0 and 1 inclusive. The same two names get the same score each time, with no network call and no memory between calls. The measure goes through one grid shared by each matcher, from 0.50 to 1.00 in steps of 0.01. A pair counts as an alert when its score is at or above the threshold, so at 1.00 a matcher that returns exactly 1 still raises an alert.

A score below 0 or above 1 is rejected and the error names the matcher. It is not forced into range and counted. The grid is computed once and rounded to the nearest whole unit, so two runs that rebuild it get the same numbers.

Check the grid

The grid and the score check are in the source, and the test suite covers both.

run it yourself
npm testthe matcher tests check the grid and the limits on the score

Where it livessrc/matcher.ts:52 · src/matcher.ts:55 · src/measure.ts:106