CRUSETRA

See what each way of comparing names costs you, live.

A threshold is the score above which two names count as a match. The figures here come from our public test set: name pairs an AI agent wrote for this repository, including near-matches that look similar but are different people. Nothing on this page comes from a customer, and the figures are recomputed as the page loads.

crusetra screening · our public test set, live

$ crusetra screen --live

each point is one matcher at one threshold, from our public test set. Pull the floor line or the slider, and the matchers that stay above it, confidence interval included, turn green

Pick a cell: each shows what it catches on confirmed matches, over its false alerts on near-matches, at that threshold
matcher \ threshold0.500.600.700.800.850.900.951.00
exact
tokens
jaro-winkler
damerau
phonetic
ngrams
embed

pick a cell: what it catches over what it flags incorrectly, with the number of pairs and the interval

$ crusetra optimise --recall

$ crusetra verify --sealed

checking…

What this rests on, for your IT auditor
  • Our public test set. releve-public.json in the repository, with a content hash (a checksum) of 556bb39ec055a091, measured at commit 318916d on 2026-09-07. The script that builds it checks that hash, rebuilds every figure with the tool’s own code, and publishes nothing if one of them disagrees.
  • The written pairs. The labeled half was written for this repository by an AI agent (commit 474ffbd): true matches, where the same name is transliterated, reordered, abbreviated or mistyped, and near-matches, which are siblings and partial homonyms that are not the same person. Where a label is debatable, the reason is written beside it.
  • The generated half. Generated from list-entry names, one kind of change at a time, and counted on their own, because a generated one is not as hard: the toggle above switches the whole grid.
  • How the tool picks. The slider sets the share of true matches you want caught. A setting counts only if the low end of its confidence interval clears that share, because a rate measured on few cases can flatter. The tool then takes the setting with the fewest false alerts.
What this page cannot do
  • Your data. This page cannot read it: no network request leaves it (the browser’s own security rules forbid them), nothing is loaded from elsewhere, and there is no input field to paste a name into.
  • It cannot show a rate without the number of pairs behind it. Each cell carries that number and its 95% confidence interval. Each matcher the tool ships with appears in our test set, and if one were ever missing it would show as a labeled blank.