Skip to content

Evaluation

Deposit → withdrawal pairs with a known answer are not publicly available. This document combines four checks, each with its own limits:

  1. a synthetic benchmark (tools/simulate.py): generated data with a known exit, to see what each component contributes and how the method fails;
  2. a placebo test on real depositors (tools/placebo_eval.py, tools/placebo_windows.py): the real pipeline run on decoy windows that cannot hold the wallet's notes, which estimates the share of chance leads per evidence family without any labels;
  3. two public laundering cases (KuCoin, Harmony);
  4. an ENS-labelled set, also scored under the protocol of Wang et al. (2023).

The placebo test is the only one that measures the signals on real withdrawals at scale. It found amount+timing and gas-price leads as frequent in decoy windows as in real ones; above chance stood the linked-address family (a direct counterparty, a withdrawal sent by the depositor's side, a deposit address swept to a labelled exchange) and, measured later, an early multi-pool profile match. A band needs one of these (since version 2.13 for the linked family, since 2.15 for the profile). Since version 2.16 strong needs lead signals from two independent sources (direct link, shared deposit address, early profile); one source is moderate, and amount+timing and gas price are context that never raises the band. Before 2.16, a lead signal plus any other family was strong, so a direct link plus a count match was strong. The synthetic tables below use the current rule. The gas-price signal turned out to fire almost only on withdrawals whose gas price a relayer chose; since 2.15 it counts only on withdrawals the user sent.

Synthetic benchmark

Setup

A fake explorer replays a depositor's deposits and the pools' Withdrawal logs through the real pipeline (multi.correlate, run_demix, ranked_candidates). Each trial draws:

  • the depositor: 2, 3 or 5 notes in the 1 ETH pool, and in half of the trials also 2–4 notes in the 0.1 ETH pool, deposited minutes apart with a hand-set gas price;
  • its exit: one address receiving every note inside a 24-hour exit window; the exit self-relays with probability 0.5, reuses the deposit gas price once with probability 0.3 and is a direct counterparty of the depositor with probability 0.2;
  • the field: N unrelated recipients per pool window (N = 10, 50, 200), each with a geometric number of withdrawals (continue with probability 0.4, so most receive one), 10 % of them active in every pool, 15 % of withdrawals self-relayed, and each withdrawal reusing the depositor's gas price by chance with probability 0.002.

Blocks are treated as pre-EIP-1559, so the gas-price signal is live, and no counterparty is a contract. 200 trials per setting, seed 1:

python tools/simulate.py --experiment all --trials 200 --seed 1

Metrics

Every recipient of a searched pool window is one instance; the depositor's exits are the positives. A listed candidate is a positive prediction.

  • P, R, F1, FPR — precision, recall, F1 and false-positive rate of "listed";
  • P ≥mod, R ≥mod — the same when only moderate and strong count as positive;
  • Top-1, Top-5 — share of trials with a true exit in the first 1 or 5 rows;
  • PR-AUC — average precision of the band-then-score order over all instances;
  • found / no lead — share of trials where a true exit is listed / nothing is listed;
  • bystander strong / moderate / weak — share of trials where the strongest band given to an unrelated address is that band.

Configurations add one component at a time: A voucher-sized counts with no gate, in no order; B the discrimination gate and a disc-scaled score; C self-relay; D gas price; E linked address; F denomination profile (the full model). G is the window-overlap suppression in the operator graph, measured separately on pairs of wallets.

Results

Ablation, 10 unrelated recipients per pool window

Config Components P R F1 FPR P >=mod R >=mod Top-1 Top-5 PR-AUC
A amount+timing, no gate 0.430 0.969 0.596 0.129 0.000 0.000 0.555 0.985 0.417
B + discrimination gate 0.433 0.966 0.598 0.127 0.000 0.000 0.685 0.995 0.722
C + self-relay 0.433 0.966 0.598 0.127 0.000 0.000 0.680 0.995 0.724
D + gas price 0.433 0.966 0.597 0.128 0.000 0.000 0.705 0.995 0.743
E + linked address 0.434 0.973 0.600 0.128 1.000 0.185 0.775 0.995 0.795
F + denomination profile 0.435 0.976 0.602 0.128 1.000 0.185 0.845 0.995 0.913

Ablation, 50 unrelated recipients per pool window

Config Components P R F1 FPR P >=mod R >=mod Top-1 Top-5 PR-AUC
A amount+timing, no gate 0.130 0.976 0.230 0.131 0.000 0.000 0.250 0.705 0.127
B + discrimination gate 0.130 0.976 0.230 0.131 0.000 0.000 0.415 0.785 0.392
C + self-relay 0.130 0.976 0.230 0.131 0.000 0.000 0.410 0.830 0.395
D + gas price 0.130 0.976 0.229 0.132 0.000 0.000 0.470 0.845 0.444
E + linked address 0.130 0.983 0.230 0.132 1.000 0.182 0.605 0.880 0.590
F + denomination profile 0.130 0.983 0.230 0.132 1.000 0.182 0.725 0.920 0.778

Ablation, 200 unrelated recipients per pool window

Config Components P R F1 FPR P >=mod R >=mod Top-1 Top-5 PR-AUC
A amount+timing, no gate 0.040 0.967 0.076 0.118 0.000 0.000 0.100 0.350 0.038
B + discrimination gate 0.040 0.967 0.076 0.118 0.000 0.000 0.130 0.490 0.156
C + self-relay 0.040 0.967 0.076 0.118 0.000 0.000 0.145 0.510 0.156
D + gas price 0.040 0.970 0.076 0.118 0.000 0.000 0.265 0.590 0.233
E + linked address 0.040 0.977 0.077 0.118 1.000 0.200 0.405 0.665 0.425
F + denomination profile 0.040 0.977 0.077 0.118 1.000 0.200 0.570 0.805 0.663

Negative controls, 50 unrelated recipients per pool window

control setting found top-1 no lead bystander strong bystander moderate bystander weak
exit after the window config A 0.00 0.00 0.06 0.00 0.00 0.94
exit after the window config B 0.00 0.00 0.06 0.00 0.00 0.94
exit after the window config F 0.00 0.00 0.06 0.00 0.00 0.94
one note per fresh address config A 0.00 0.00 0.10 0.00 0.00 0.90
one note per fresh address config B 0.00 0.00 0.10 0.00 0.00 0.90
one note per fresh address config F 0.00 0.00 0.10 0.00 0.00 0.91

Counter-measures against the full model, 50 unrelated recipients

counter-measure setting found top-1 no lead bystander strong bystander moderate bystander weak
delay 0 of 3 notes after the window 1.00 0.17 0.00 0.00 0.00 0.99
delay 1 of 3 notes after the window 0.00 0.00 0.01 0.00 0.00 0.99
delay 2 of 3 notes after the window 0.00 0.00 0.01 0.00 0.00 0.99
delay 3 of 3 notes after the window 0.00 0.00 0.01 0.00 0.00 0.99
fresh addresses 1 exit address(es) 1.00 0.17 0.00 0.00 0.00 0.99
fresh addresses 2 exit address(es) 0.00 0.00 0.01 0.00 0.00 0.99
fresh addresses 3 exit address(es) 0.00 0.00 0.00 0.00 0.00 1.00
relayer self-relayed 1.00 0.17 0.00 0.00 0.00 0.99
relayer through a relayer 1.00 0.17 0.00 0.00 0.00 0.99
gas strategy deposit gas price reused 1.00 0.17 0.00 0.00 0.00 0.99
gas strategy wallet default gas 1.00 0.17 0.00 0.00 0.00 0.99
denominations both pools to one exit 1.00 0.86 0.00 0.00 0.00 1.00
denominations each pool to its own exit 1.00 0.18 0.00 0.00 0.00 1.00

Counter-measure intensity, six-note voucher (full model)

python tools/simulate.py --experiment intensity: the share of trials in which the exit is listed, against the number of notes withdrawn after the window and the number of exit addresses the notes are spread over.

Notes withdrawn after the window (of 6):

Exit leaves 0 1 2 3 4 5 6
amount and timing only 1.00 0.00 0.00 0.00 0.00 0.00 0.00
also a direct transfer 1.00 1.00 1.00 1.00 1.00 1.00 0.00

Exit addresses the six notes are spread over:

Exit leaves 1 2 3 4 5 6
amount and timing only 1.00 0.00 0.00 0.00 0.00 0.00
also a direct transfer 1.00 1.00 1.00 1.00 1.00 1.00
also a direct transfer, share of exits found 1.00 0.50 0.33 0.25 0.20 0.17

One delayed note or a second exit address is enough to remove an exit that left only amount and timing. An exit that also transacted with the depositor is found as long as one note lands in the window; with k exit addresses only that one is found, so the share of exits found falls as 1/k. No bystander reaches strong or moderate at any intensity: without a linked address a chance count or gas-price match stays weak.

Wallets merged without suppression merged with suppression
unrelated, identical fingerprints, same hour 1.00 0.00
unrelated, distinct fingerprints, same hour 1.00 0.00
one operator, distinct fingerprints, days apart 0.94 0.91

Sensitivity to the field assumptions (full model)

Field Top-1 PR-AUC control: strong control: moderate control: no lead
baseline 0.72 0.78 0.00 0.00 0.06
no chance gas reuse 0.77 0.87 0.00 0.00 0.07
chance gas reuse x5 (p=0.01) 0.67 0.79 0.00 0.00 0.09
self-relay 5 % 0.74 0.79 0.00 0.00 0.06
self-relay 30 % 0.70 0.76 0.00 0.00 0.06
busier recipients (p_more=0.6) 0.66 0.77 0.00 0.00 0.01

Reading the results

  • The gate and the score order the field; they do not shrink it. A and B list the same addresses at these field sizes (every voucher-sized count in a field of 10+ passes disc ≥ 0.5), but B's disc-scaled score lifts PR-AUC about threefold (0.13 → 0.39 at N = 50).
  • On generated data each family adds ranking power. Through B, D, E and F, PR-AUC rises 0.39 → 0.44 → 0.59 → 0.78 and top-1 0.42 → 0.73 at N = 50; at N = 200 top-1 rises from 0.13 to 0.57. The gas-price step is smaller than before version 2.15: a generated exit that reuses the deposit gas price through a relayer no longer counts (in the counter-measure table, top-1 for "deposit gas price reused" falls from 0.99 to 0.17), which is what the placebo test says about real relayed withdrawals. Recall of "listed" stays near 0.98: the exit is almost always among the candidates, and the question is how high it ranks. The placebo test below shows that the amount+timing and gas-price part of this does not carry over to real pools.
  • A band now means a linked address. moderate is reached only by exits that transacted with the depositor (20 % of generated exits), so P ≥mod is 1.00 and R ≥mod about 0.18-0.20. The generated field has no shared deposit address and no early profile, so strong (two lead sources) never arises in it; the tables are byte-identical under the 2.16 rule. In the negative controls, where no exit is findable, no unrelated address reaches moderate or strong (before version 2.13: moderate in about 80 % and strong in 11 % of trials, through self-relay and chance gas-price reuse).
  • Self-relay no longer lowers PR-AUC (0.392 → 0.395 at N = 50): it used to lift self-relayed bystanders to moderate; it is now context inside the amount+timing family.
  • Counter-measures remove the exit, not the candidates. Delaying one note past the window or splitting the notes over fresh addresses drops "found" from 1.00 to 0.00; the tool still lists unrelated addresses, now all weak, and "no lead" stays near 0.
  • Suppression removes false operator links. Unrelated wallets that deposit in the same hour are merged in every trial by a naive shared-candidate rule and in none with the window-overlap suppression, while two wallets of one operator that share a consolidator are still merged in 91 % of trials (94 % naive).

Placebo test on real depositors

tools/placebo_eval.py draws random Ethereum depositors from the ENS universe (deposits spanning at most 60 days, windows complete before August 2026; seed 1) and runs the real pipeline twice per depositor: on the usual windows after its deposits (target), and on the same deposits shifted back so that every window ends a day before its first real deposit (decoy). A withdrawal in a decoy window cannot spend one of the wallet's notes, so every decoy lead is a false note link; everything else — history, counterparties, deposit gas prices — is the wallet's own. The share of target leads that chance explains (the ratio R below) is estimated as decoy leads per withdrawal searched over target leads per withdrawal searched; 95 % intervals come from a bootstrap over depositors. tools/placebo_windows.py re-scores the cached runs offline for narrower exit windows (it reproduces the 30-day bands exactly in all 304 runs).

152 depositors, 30-day window, leads that reached strong or moderate under the previous band rule (so that every family is visible), by family:

Family Target leads Decoy leads Chance share (95 % CI)
linked address (linked, linked_sender) 30 7 0.25 (0.07-0.54)
amount+timing (count match, self-relay) 160 188 1.25 (0.79-1.81)
gas price 45 64 1.51 (0.92-2.31)
all, previous rule 228 252 1.17 (0.82-1.60)

(268,745 withdrawals searched in target windows, 253,634 in decoy windows.)

The same families by exit window (target / decoy leads):

Window linked address amount+timing gas price
6 h 14 / 0 3 / 2 1 / 1
24 h 16 / 0 11 / 12 8 / 2
72 h 19 / 1 (0.06, 0-0.24) 20 / 20 12 / 7
30 days 30 / 7 (0.25, 0.07-0.54) 160 / 188 45 / 64

Of the 152 depositors, 115 deposited after EIP-1559: for them the gas-price signal never fires (the base-fee gate) and amount+timing has a chance share of 1.2-1.9 at every window. The 37 earlier depositors give too few narrow-window leads to decide either family.

What this shows:

  • Amount+timing and gas price do not separate real exits from chance on real depositors, at any window from 6 hours to 30 days. The synthetic benchmark says they rank a planted exit well; on real pools the same count or gas price turns up as often where no note of the wallet can be.
  • The linked-address family does. Its leads are four times as frequent in real windows as in decoy ones over 30 days, and within 72 hours of the deposit 19 against 1. That is why a band needs a lead signal such as a linked address, and why a linked exit within 72 hours is marked as an early exit. Since 2.17 a direct link counts only within 72 hours (below).
  • A decoy lead is a false note link, not necessarily a false identity link: a linked address in a decoy window may still belong to the depositor. For the linked family the chance share is therefore an upper bound.
  • The sample is 152 random depositors, most of whom leave no linked exit at all; the intervals are wide, and the result says nothing about careful users beyond the fact that the tool finds no evidence on them.

Why gas price failed. Of the gas-price matches in these runs, 44 of 45 in target windows and 62 of 64 in decoy windows were on withdrawals sent through a relayer: the relayer chose that gas price, not the user, so a match with the deposit's gas price says nothing about who withdrew. Since version 2.15 the signal counts only on withdrawals the user sent (no relayer); re-scored with that rule, the same 152 runs give no gas-price lead at all. The tables above keep the previous rule, under which the test was run.

Early multi-pool profile

tools/placebo_profile.py tests a denomination-profile match offline on every eligible depositor of the ENS universe (28,739 depositors with two or more notes; no explorer calls). The profile is the wallet's note count per pool; a recipient matches when, in every pool the wallet used, it received exactly that many withdrawals between the pool's first deposit and its last deposit plus a window. Decoy windows are shifted back as above.

Profile 24 h 72 h 30 days
2+ notes, any pools 117724 / 108163 (0.99) 175944 / 164315 (1.00) 911742 / 876567 (1.01)
2+ notes, 2+ pools 3222 / 1282 (0.43) 6509 / 4275 (0.71) 57693 / 52887 (0.96)
6+ notes, any pools 3551 / 2394 (0.74) 4809 / 3357 (0.77) 19681 / 16874 (0.90)
6+ notes, 2+ pools 784 / 114 (0.16) 1089 / 331 (0.34) 3992 / 2966 (0.78)
10+ notes, any pools 477 / 206 (0.47) 605 / 265 (0.48) 1557 / 1187 (0.81)
10+ notes, 2+ pools 248 / 13 (0.06) 318 / 34 (0.12) 642 / 314 (0.52)

(target / decoy hits, chance share.) A single-pool profile is the count match again and stays near chance. A profile over two or more pools is far rarer by chance, and the shorter the window the cleaner: with at least 10 notes over two or more pools (4,194 depositors), 318 hits against 34 within 72 hours (chance share 0.12, 95 % interval 0.08-0.16). A direct linked address is cleaner in the same 72-hour window (0.06), but the profile rests on different data than an address link. Since version 2.15 this is the early_profile signal (at least 10 notes, two or more pools, 72 hours); it makes a lead on its own. Being a stricter count match, it belongs to the amount+timing family; as one lead source it is moderate, and it reaches strong only together with a second lead source (a direct link or a shared deposit address).

Hold-out and sensitivity. tools/placebo_profile.py takes --period before|after (split at the sanctions, 2022-08-08), --notes, --hours and --out; its default output is unchanged. The rule as used by the tool (10+ notes, 2+ pools, 72 h) on each half:

Period Depositors Target / decoy hits Chance share (95 % CI)
before 2022-08-08 1,448 142 / 5 0.04 (0.01-0.08)
after 2022-08-08 2,746 176 / 29 0.19 (0.12-0.27)
whole set 4,194 318 / 34 0.12

The rule is cleaner before the sanctions than after, but it holds in both halves. Sensitivity on the whole set (2+ pools, 72 h): 6+ notes 1,089 / 331 (0.34), 10+ notes 318 / 34 (0.12), 15+ notes 112 / 7 (0.07); a 168 h window with 10+ notes gives 383 / 79 (0.22). The thresholds were chosen from the whole-set table above and are checked here on each half, so this is a consistency check rather than a fully independent hold-out; a lower note threshold admits many more chance matches, a higher one finds fewer exits.

Permutation control. tools/placebo_profile_perm.py (output profile_perm.json) keeps the real 72 h windows and replaces the depositor's profile with that of another depositor with the same pool set. The own profile gives 319 hits; the other depositors' profiles give a mean of 45 (min 28, max 61) over 20 rounds, a ratio of 0.14. The hits therefore come from the depositor's own note counts, not from the pool set or the window alone.

Other chains. tools/ens_labels.py universe --network ... --extra ... builds the same deposit/withdrawal universe for another chain (on L2 chains deposits go through the Tornado proxy 0x0D5550d52428E7e3175bfc9550207e4ad3859b17, which the registry does not list), and tools/placebo_profile.py --universe and tools/placebo_profile_perm.py --universe read it. BSC, Gnosis and Base are not served by the free explorer tier, so Arbitrum (the same four ETH pools) and Polygon (100, 1,000 and 10,000 MATIC) were used; universes from June 2021 to September 2026. The rule as used by the tool (10+ notes, 2+ pools, 72 h):

Chain Deposits / withdrawals Eligible depositors Target / decoy hits Chance share (95 % CI) Permutation: own / borrowed profile
Ethereum (see above) 4,194 318 / 34 0.12 (0.08-0.16) 319 / 45
Arbitrum 6,721 / 6,315 74 16 / 0 0 17 / 0.2
Polygon 20,194 / 19,428 253 50 / 1 0.02 (0.00-0.09) 54 / 4.9

The early profile stands above chance on both chains; the samples are small, Arbitrum's especially. The linked-address placebo test was not repeated there.

Does a second family make a lead stronger?

Before 2.16, a lead signal plus any other family was strong. tools/placebo_strong.py (output strong.json) checks this on the 152 depositor runs: among candidates with a lead signal, how often does a second family appear in real versus decoy windows?

Candidates with a lead signal and 72 h target / decoy 30 days target / decoy
amount+timing 4 / 1 (chance share 0.30, 95 % CI 0-1.87) 4 / 0
nothing else 17 / 0 28 / 7 (0.27)

A count match next to a lead did not lower the chance share, and the counts are small. So amount+timing and gas price no longer raise the band; they are shown as context. Under the 2.16 rule these 152 runs give no strong lead: they predate shared_deposit, so strong could only have come from a direct link plus an early profile. moderate (one lead source) is 21 / 1 at 72 h (0.06, CI 0-0.23) and 32 / 7 at 30 days (0.23, CI 0.07-0.53).

The 152 runs are mostly from before the sanctions. A second sample of 150 depositors whose first deposit is on or after 2022-08-08 (tools/placebo_eval.py --after 2022-08-08 --dir post2022 --seed 3) was run the same way, and tools/placebo_periods.py splits both by period. Together they hold 302 depositors (210 after the sanctions: 150 plus 60 of the first sample). Splitting the direct-link leads (linked, at the 30-day window) by the delay of the recipient's first withdrawal after the voucher's last deposit (target / decoy leads):

First deposit within 72 h later than 72 h
before 2022-08-08 15 / 0 5 / 4
on or after 2022-08-08 11 / 1 3 / 9
all 302 26 / 1 8 / 13

A direct link withdrawn within 72 hours stands well above chance in both periods; a later one is at chance level. For the 210 post-2022 depositors the whole linked-address family gives 12 / 1 within 72 hours (chance share 0.11) but 14 / 12 within 30 days (0.95). The pooled number above, 30 against 7 over 30 days, hid this: it mixed early links, which separate, with late ones, which do not. Since 2.17 a direct link is linked only when the first withdrawal came within 72 hours (EARLY_EXIT_HOURS); a later one is linked_late, listed in the evidence with its delay, never scored and never a lead. The withdrawal-sender signal gave 3 / 0 early and 5 / 2 late, too few to decide, and in the KuCoin case the attacker's real exits started 6.9 days after the deposit, so linked_sender is left as is; so are shared_deposit and early_profile. The tables above keep the rule under which they were measured.

Re-scored under 2.17 (all strong and moderate leads, real / decoy):

Sample within 72 h within 30 days
152 depositors (first sample) 21 / 1 (0.06, 0-0.21) 25 / 1 (0.04, 0-0.17)
210 post-2022 depositors 12 / 1 (0.11, 0-0.53) 12 / 3 (0.28, 0-0.90)
all 302 depositors 31 / 1 (0.04, 0-0.15) 35 / 3 (0.09, 0-0.27)

(Before 2.17 the first sample gave 32 / 7, 0.23, over 30 days.) On the ENS set 2.17 finds 12 of the 31 pairs (21 inside the window) as moderate or strong among 28 such candidates (precision lower bound 43 %, against 16 of 40 before); the two strong candidates are both labelled pairs; four labelled pairs fall to weak because their direct link came later than 72 hours. Under the protocol of Wang et al. demix up to moderate scores precision 1.00, recall 0.41, F1 0.59 (0.71 before). The evaluation tools re-apply the current heuristics to cached runs, so these numbers follow the rules as shipped.

Bybit (2025)

The Bybit theft of 21 February 2025 is a later, independently labelled case: the repository of Liu et al. (Evasion Under Blockchain Sanctions, WWW 2026) lists about 9,300 tracked and 12,200 Elliptic-flagged Ethereum addresses. Nine of them deposited into the Ethereum ETH pools (24 February to 2 April 2025; 1 to 22 notes). demix (30-day window, attribution set loaded) gives one moderate lead in total, a withdrawal back to the depositor's own address, and one listed address among the weak candidates (a count match). No early multi-pool profile fired for the three depositors with 11 to 22 notes, and no unlisted address reached moderate or strong. The launderers spread the exits, so the method finds almost nothing here, but it raises no false lead either.

Signals tested and not used

Priority tip after EIP-1559. Since London a sender's gas price is the block's base fee plus a priority tip, so the old gas-price signal cannot fire; the tip is the part a user chooses. tools/placebo_tip.py computes tips for every self-relayed withdrawal and every deposit of the 302 placebo depositors (base fees from a public node) and flags a recipient whose tip equals one of the depositor's deposit tips and is rare in the pool window:

Rarity rule 72 h, target / decoy 30 days, target / decoy
tip shared by at most 3 self-relayed withdrawals 33 / 44 (1.57, 0.98-2.58) 60 / 70 (1.26, 0.83-1.97)
at most 10 102 / 105 (1.21) 181 / 204 (1.22)

Tips cluster on wallet defaults (3, 0.5, 1 and 2 gwei), so a shared tip is chance-level and the signal is not used. Wallet fingerprints (gas limit, transaction type, fee settings) belong to the same class of client defaults and would need a transaction lookup per withdrawal; given this result they were not pursued. Anonymity-mining reveals (AP/TORN claims) apply only to 2020-2021 deposits, too few in these samples to test.

Does the discrimination gate remove chance matches?

tools/placebo_disc.py (output disc.json) bins count matches (every recipient whose withdrawal count equals a voucher size) by discrimination, in real and decoy windows:

Bin Window Target / decoy Chance share (95 % CI)
all count matches 72 h 8,966 / 7,640 1.01
gated in (disc ≥ 0.5) 72 h 717 / 591 0.98 (0.80-1.20)
disc ≥ 0.95 72 h 119 / 117 1.17
gated in 30 days 5,569 / 5,256 1.00

The gate cuts the number of count matches about twelvefold, but the chance share stays near 1 in every bin: it reduces the field the analyst has to read and does not separate real count matches from chance ones.

Shared exchange deposit addresses

The shared-deposit-address signal (shared_deposit, see METHODOLOGY.md) was tested the same way. tools/placebo_dar.py applies it to the 152 depositors above; tools/placebo_dar_universe.py reads the windows offline from the ENS universe (it reproduces the full-run result for all 152 depositors) and draws 1,000 new random depositors (seed 2). Lookups that hit an explorer error are retried rather than silently skipped; none failed in the final runs. Version 2.15 code, attribution set loaded unless stated.

Sample and rule Depositors with a deposit address Target hits Decoy hits Chance share (95 % CI)
152 depositors (full runs): swept to a labelled exchange (scored) 4 0 0
152 depositors: any deposit address found 28 9 1 0.12 (0-0.53)
1,000 depositors: swept to a labelled exchange (scored) 27 (19 depositors) 4 0.16 (0.03-0.44)
1,000 depositors: any deposit address found 251 38 (28 depositors) 13 0.36 (0.11-0.89)
1,000 depositors, no attribution set (activity only) 164 16 (13 depositors) 12 0.80 (0.21-2.37)

(About 2.1 million withdrawals searched in target windows and 2.0 million in decoy windows.) An address counts as a deposit address when its outflow goes to hot wallets. A hot wallet recognised by an exchange label gives a clean signal: 27 against 4, and 25 of the 27 are not direct counterparties of the depositor, so the signal finds exits that linked does not. A hot wallet recognised by activity alone does not: without an attribution set the signal is at chance (16 against 12). Two fixes on the way did not change that: a busy contract (a token, a DEX router, the Tornado router) is no longer taken for a hot wallet, and an unlabelled hot wallet must receive amounts forwarded within 3,200 blocks (the forwarding test of Victor used by Tutela). So since version 2.15 a shared deposit address is scored only when it sweeps to labelled exchange wallets; one found by activity alone is shown as context. Without an attribution set the signal therefore does not fire.

(History: 2.14.0 published 57 against 18 for "label or activity"; that run used a busy check the explorer's page cap had turned off, and later ones counted token and router contracts as hot wallets. The "labelled exchange" numbers are the same in every run.)

The check is also available per case (demix --placebo, and the web UI option): the decoy window of the wallet under investigation, beside the real one.

Current rule (2.17) in one table

tools/review4_*.py re-score the cached runs offline under the 2.17 rule. R is decoy leads per withdrawal searched over target leads per withdrawal searched; below 1 it estimates the chance share of target leads, near or above 1 the signal is chance level. Intervals: bootstrap over depositors, or exact (*) when a side has fewer than five leads.

Signal or class Depositors Window Target / decoy R (95 % CI)
early direct link (lead) 302 30 d 26 / 1 0.04 (0.00–0.16)
late direct link (context) 302 30 d 8 / 13 1.76 (0.69–6.4)
linked sender (lead) 302 30 d 8 / 2 0.27 (0.03–1.36)*
shared deposit, labelled exchange (lead) 1,000 30 d 27 / 4 0.16 (0.03–0.44)
shared deposit, no labels (context) 1,000 30 d 16 / 12 0.80 (0.21–2.37)
early multi-pool profile (lead) 4,194 72 h 318 / 34 0.12 (0.08–0.16)
count match (context) 302 30 d 9,305 / 8,540 0.99 (0.92–1.07)
gas price (context) 302 30 d 1 / 2 too few
moderate 302 72 h 30 / 1 0.04 (0.00–0.15)
moderate 302 30 d 34 / 3 0.10 (0.00–0.29)
strong 302 30 d 1 / 0 too few
strong + moderate 302 72 h / 30 d 31 / 1, 35 / 3 0.04, 0.09

The first 152 runs predate the shared-deposit signal; adding its later-computed labelled hits gives 33 / 1 and 38 / 3 for strong + moderate and 2 / 0 for strong.

strong is structural, not measured. It appeared only in target windows (one or two cases, each an early direct link plus a labelled shared deposit — the same combination as both strong candidates of the ENS set). On the 1,000 depositors an early profile never coincided with a labelled shared deposit in either window; a direct link plus shared deposit cannot be counted there (no counterparty history was fetched). Only moderate has a measured chance share, so the tool suggests the same independent check for both classes.

The 72 h split of the direct link was decided on these data (the 72 h bound itself came from the earlier early-exit flag) and the 150 post-2022 depositors were part of that analysis, so the 2.17 rule is exploratory until checked on new depositors. The early profile differs between periods (0.04 before the 2022 sanctions, 0.19 after; the intervals do not overlap) and fires for 4,194 of the 64,681 depositors in the universe (about 6.5 %).

Two public laundering cases

tools/real_cases.py runs the tool on the depositors of two cases attributed to the Lazarus group, with addresses from the public investigations collected at github.com/tayvano/lazarus-bluenoroff-research. Default settings, 30-day window.

Case Depositors (deposits) Exits known Exits in a window Candidates True Other
KuCoin 2020 2 (55) 36 35 per depositor 6 6 0
Harmony 2022 14 (857) 30 of 55 listed 0-30 per depositor 0 0 0

Run with python tools/real_cases.py; it also prints the context marks below.

KuCoin. Koh (2020, "Deanonymising the Kucoin Hacker") attributes to the attacker two addresses that called withdraw() themselves without a relayer: 0x8bd8... (323 withdrawals) and 0x82e6... (128, of which 114 straight to the 100 ETH pool and 14 through the router); together they sent the 437 withdrawals from the 100 ETH pool counted there. Their recipients, 36 addresses, are the ground truth; no demix signal uses the transaction sender, except linked_sender when the sender is a counterparty of the depositor. The check covers the two case addresses that deposited themselves (55 notes); Koh counts 497 deposits of 100 ETH for the attacker overall, so this is a slice of the case.

The count match finds nothing: exits received many notes each, none equal to a voucher (24 or 30 notes), because they collected notes of several deposits. With linked_sender, the depositor 0x820a..., which had transacted with 0x82e6..., gets 6 candidates, all true exits of that caller, all moderate; the first of these withdrawals came 6.9 days after its last deposit. The 30 exits of 0x8bd8..., which started withdrawing a day after that deposit, are not found: 0x8bd8... is not a counterparty of either depositor. So of the attacker's 35 exits inside the window the tool finds 6. How many of the others carry notes of these two depositors (55 of the attacker's 497 notes) cannot be told from chain data, so this is not a per-depositor recall: the second caller's notes may come from other deposits. The other depositor used to get one unrelated candidate from a chance gas-price match on a relayed withdrawal (moderate before version 2.13, weak in 2.13-2.14); since 2.15 a relayed withdrawal earns no gas-price signal and the candidate is gone.

Harmony. Investigators listed 55 withdrawal addresses; 30 of them received 180 withdrawals from the 100 ETH pool within 30 days of the deposits, 6 per address in most cases, against vouchers of 60 notes. The tool lists no candidate and no false one. The investigators selected these addresses partly by withdrawal count and batching, so this list is not independent of count-based reasoning.

Context marks. Every true exit that fell in a window had at most a day of history before its first withdrawal: 35 of 35 in KuCoin, 30 of 30 in Harmony. The fresh mark is shown as context and never scored, because a new unrelated wallet is fresh as well. The shared-funder edge links the two KuCoin depositors through 0x0060..., an address on the investigators' list. In Harmony it groups all 14 depositors into 5 clusters through 5 funders, each shared by 2-3 depositors; each of the 5 funders received its funds directly from the bridge exploiter 0x0d04...ded00. A sixth funder, shared by 4 depositors, has 200 or more transactions and is left out as busy.

Exit groups. In Harmony, 287 of the 295 (true exit, depositor window) pairs sit in an exit group. With one true exit as the only anchor, the group holds 23.0 addresses on average, 16.8 of them (73 %) on the investigators' list; the other members were paid out in the same bursts and are leads, not errors, as the list itself is partial. In KuCoin, 30 of the 70 (true exit, depositor window) pairs sit in a group; from one true exit the group holds 4.9 addresses, 4.5 of them (91 %) true exits.

Both cases show the same limit as the synthetic counter-measures: when notes are pooled and redistributed, a count match finds nothing, and the tool reports no lead rather than a wrong one. Real laundering elsewhere shows the same pattern: in the 27 MixLaunder cases, direct linkage covers 1.27 % of laundering addresses (4 cases), the same address on both sides 1.48 % (3 cases).

Limits

The field model is simple: recipients are independent, gas prices are drawn from a short list of defaults, exits never reuse addresses across cases, and every exit lands inside the window unless a counter-measure says otherwise. Real pools have bursts, relayer-specific patterns and post-EIP-1559 fee markets. The benchmark is a controlled test of the method's components and failure modes; on real depositors the placebo test below shows that its amount+timing and gas-price results do not carry over.

A labelled set from ENS

There is no public set of matched deposits and withdrawals. As in Béres et al. (2021), Tutela (2022) and Wang et al. (2023), ENS provides a partial one: tools/ens_labels.py collects every deposit into and withdrawal from the four Ethereum ETH pools from December 2019 to September 2026 (289,887 deposits by 64,681 depositors; 276,676 withdrawals to 129,752 recipients), reads the primary ENS name of every address (4,606 have one) and links two addresses when one controls the other's name (447 links; an owner of more than five names is taken for a service and ignored). A depositor and a recipient of the same pool linked this way, with the withdrawal after the deposit, form a labelled pair: 31 pairs of 27 depositors. Only 2 of them deposited after the sanctions of August 2022.

tools/evaluate_labels.py runs demix with default settings (30-day window, attribution labels loaded) on every labelled depositor:

strong up to moderate all bands
Pairs found (of 31; 21 inside the window) 2 12 (57 % of those in the window) 16
Precision, lower bound 2 of 2 12 of 28 (43 %) 16 of 2,811
Pairs found without the linked signal 0 2 3

(Version 2.17. Four labelled pairs fall to weak because their direct link came later than 72 hours; under 2.16 the column read 16 found, 16 of 40 (40 %). Under the band rule before 2.13, where a gas-price match or a self-relayed count match alone reached moderate, the lower bound was 16 of 85, 19 %. Before 2.16, strong was a lead signal plus any other family: 1 pair found, 1 of 4 candidates.)

All 16 pairs are found through a direct transaction between the depositor and the recipient (linked). The ENS link and linked see the same relationship, so this confirms that careless users also transact directly; it does not measure the other signals. Two of the 16 pairs are also found by shared_deposit (depositor and recipient sent funds to the same deposit address of a labelled exchange); with the direct link that is two lead sources, so these two are the strong pairs, and without linked they stay moderate. Otherwise the other signals find almost nothing here, for a reason the method states in advance: 24 of the 31 pairs come from depositors whose only voucher in that pool is a single note, where the count match cannot narrow the field (discrimination 0.12-0.33 against the 0.5 threshold), and a gas-price match is rarely checkable after EIP-1559. The count match fires for one pair only (a two-note voucher); with a direct link that is one lead source, so that pair is moderate (before 2.16 it was strong). Ten pairs fall outside the window (withdrawals 50-1,794 days after the deposit).

The precision is a lower bound. The labels say nothing about the 24 other strong and moderate candidates, all direct counterparties of the depositor (3 are the depositor itself; 16 withdrew within 72 hours of the deposit). The set covers only users careless enough to put an ENS name on both sides, names are read as they are today, and a pair links two addresses, not a deposit to a withdrawal. The pairs link named people to Tornado Cash use, so they stay local; only these aggregates are published.

The same set under the protocol of Wang et al.

Wang et al. (2023) validate their heuristics at the address level: test pairs are every labelled depositor times every labelled withdrawer, and a predicted pair that is not labelled counts as a false positive. Their average F1 of 0.55 comes almost entirely from H3, a direct transfer between the two addresses. tools/wang_baseline.py re-implements H2 (the depositor sent the withdrawal), H3 and H5 (cross-pool deposit profile) from the paper and applies the protocol to the 27 × 29 labelled addresses here:

Method Precision Recall F1
Wang H2 1.00 0.03 0.07
Wang H3 (direct transfer, whole history) 1.00 0.93 0.96
Wang H5 0.00 0.00 0.00
demix, up to moderate 1.00 0.41 0.59
demix without linked 1.00 0.07 0.13

A direct-transfer check alone, with no mixer analysis at all, scores 0.96, because 27 of the 29 pairs transacted directly. Labels built from address links (ENS, airdrop aggregation) measure whether two addresses ever met, not whether a withdrawal spends a deposit; this holds for the F1 of 0.55 in Wang et al. as much as for the numbers here. demix recalls less than H3 because it only looks at withdrawals in the 30-day window after the deposit.

Do the signals move together?

On the placebo runs under the current signal set (302 depositors, 30-day window: 283,383 recipient-pool rows in target windows, 254,263 in decoy windows; tools/review4_analyze.py), only one pair has enough expected co-occurrences to test: count match and self-relay (555 together against 252 expected in target windows, 526 against 263 in decoy windows; phi 0.037 and 0.033). Both read the same withdrawals and belong to one family, which is why the score takes only the larger weight. Every other pair has an expected count below one. Cross-source co-occurrences of lead signals appear only in target windows (twice: early direct link plus shared deposit) and never in decoy windows. The early profile coincides with a count match by construction (2 against 0.07), but a count match is not a lead source. Independence of the lead sources is therefore an assumption of the model, not a measurement.

(An earlier check on the ENS, KuCoin and Harmony runs, made before the shared-deposit and early-profile signals and the gas-price restriction, found every pair within three co-occurrences of independence on 43,424 rows.)

Checks of the code itself

Run on 2026-09-28; tests, coverage and the live check re-run on 2026-10-02.

Check Result
pytest (network blocked, incl. getaddrinfo) 755 passed
pytest-randomly, seeds 1, 2, 3 all pass in every order
Branch coverage (--cov-branch) 93 % overall; heuristics 97 %, demix 93 %, cli 89 %
Property-based tests (Hypothesis, 24) and edge-case tests (10) pass; the four that documented defects now pin the fixes
Mutation testing (mutmut 3.8 on heuristics.py and demix.py, 3061 mutants, version 2.17.0; the static encoding check deselected) 2305 killed (75 %), 754 survived, 2 not reached (2.16: 74 % of 2952)
pip-audit -r requirements-lock.txt no known vulnerabilities
pytest -m live (every shipped pool re-verified on chain) 56 pools on 8 networks verified (Polygon via drpc.org: publicnode returns empty eth_call results there)
README cases re-run (Ronin, Wintermute, Beanstalk) same figures: 12,595.3 ETH hop; one 9.9435 ETH inflow; 271 deposits in 2.97 h, one weak candidate

Most surviving mutants change log and evidence wording, result keys that no test reads, or fallback values that the pipeline never reaches (a signal without a weight, a boundary that no input can hit). Mutation testing raised the score from 65 % by pinning what it found in the core: transfer mode had no test at all, and the voucher span limit, counterparty collection, the shape of a voucher and the end-to-end run of run_demix were unpinned at their edges. mutmut does not run natively on Windows; the run used the python:3.13-slim container.