Model risk · validation report · aml_triage
Validation report: Alert triage for transaction monitoring
Harborline Bank is fictional. This is a demonstration on generated data and public datasets; no real customer, account or transaction of any real institution appears here.
Nothing here is a credit decision, a fraud determination, a suspicious activity finding or investment advice.
| Model | aml_triage |
| Purpose | Rank the transaction monitoring rules' alerts so an investigator opens the likeliest real typology first |
| Tier | 2: Informs a decision a person makes; moderate materiality |
| Role | sole |
| Status | in use |
| Owner | Head of analytics |
| Prediction time | When the alert fires; history features use only the account's entries before the alert window opened |
| Outcome | The alert overlaps an injected typology episode of the same kind on the same account |
| Last validated | 2026-08-31 |
| Conclusion | approved |
Purpose and scope
The transaction monitoring rules fire alerts; investigators work them in the order the queue presents them. This model sets that order: the alert most likely to be a real typology comes first. It informs which alert a person opens, never whether an alert is closed or reported, so it is tier two.
Out of scope: deciding a disposition, filing a suspicious activity report, and any alert
source other than the four rules in rules/aml_typologies.yaml.
Conceptual soundness
Triage is a ranking problem over a small, imbalanced set, so the method is a logistic regression on standardised features whose coefficients an investigator can read: the alert's own size, entry count, counterparties and branches, which rule fired, the account's age, its cash deposits and inflow in the ninety days before the alert window opened, and how many alerts it had before. The benchmarks are the base rate and the order the rules alone imply, each rule ranked by its precision on the development months.
Data and assumptions
Every alert the rules fired on Harborline Bank's deposit accounts: 649 alerts, 429 of them overlapping an injected typology episode of the same kind on the same account. The label is the generator's truth, which a real program would not have; its labels would be investigator dispositions.
Split by when the alert closed: development September 2023 to August 2025 (445 alerts), validation September 2025 to February 2026 (98), test March to August 2026 (106, 59 true).
A leakage test recomputed the history features of 200 sampled alerts from raw ledger entries strictly before each alert window and matched every one. No demographic attribute or proxy is a feature.
Development evidence
Development AUROC 0.999, validation 1.000. The coefficients, standardised, largest first:
| feature | coefficient |
|---|---|
| prior_cash_90d | -3.32 |
| branches | 1.86 |
| rule_structuring | -1.56 |
| prior_alerts | 0.94 |
| rule_round_tripping | 0.82 |
| rule_rapid_movement | 0.81 |
| counterparties | 0.60 |
| rule_mule | 0.57 |
| log_amount | 0.55 |
| entries | 0.40 |
| log_tenure_days | -0.23 |
| log_prior_inflow_90d | 0.16 |
| amount_vs_monthly_inflow | 0.03 |
Outcomes analysis on held out data
On the test months the model reaches an AUROC of 1.000 and an average precision of 1.000. Within the 72 structuring alerts, 25 of them true, where the rules alone cannot order anything, its AUROC is 1.000.
These are properties of the generator, not of the method: its legitimate lookalikes and its injected typologies do not overlap. Without the cash history and branch count the AUROC is 0.997; the rules' own order reaches 0.788.
Calibration on the test sample, ten equal count bins:
| Bin | Accounts | Predicted | Observed |
|---|---|---|---|
| 1 | 11 | 0.04% | 0.00% |
| 2 | 11 | 0.27% | 0.00% |
| 3 | 11 | 0.91% | 0.00% |
| 4 | 11 | 3.67% | 0.00% |
| 5 | 11 | 81.86% | 72.73% |
| 6 | 11 | 99.39% | 100.00% |
| 7 | 10 | 99.83% | 100.00% |
| 8 | 10 | 99.92% | 100.00% |
| 9 | 10 | 99.96% | 100.00% |
| 10 | 10 | 99.99% | 100.00% |
Source: generated, Every alert the four rules fired on Harborline deposit accounts, September 2023 to August 2026, seed 20260831, as of 2026-08-31.
Benchmarking
Against the base rate, the rules' own order, and this model with its two strongest features removed, on the test months:
| Model | AUROC | Gini | Brier | ECE |
|---|---|---|---|---|
| This model | 1.000 | 1.000 | 0.0071 | 0.0155 |
| Base rate | 0.500 | 0.000 | 0.2601 | 0.1489 |
| The rules' own order | 0.788 | 0.576 | 0.1576 | 0.1276 |
| This model without cash history and branches | 0.997 | 0.994 | 0.0362 | 0.0751 |
Source: generated, Every alert the four rules fired on Harborline deposit accounts, September 2023 to August 2026, seed 20260831, as of 2026-08-31.
Limitations
- The labels are injected truth. An investigator's disposition is noisier, arrives weeks later and is shaped by the queue order itself, which this data cannot show.
- The legitimate lookalikes are written to look like the typologies on the rule's terms, not on every feature, so the model's separation here overstates what it would do on a real portfolio.
- Four rules, four typologies. Anything the rules do not fire on never reaches the queue.
Ongoing monitoring plan
Monthly, against the queue as investigators work it:
| Measure | Trigger | Action |
|---|---|---|
| Precision of the queue's first quarter against investigator dispositions | below the rules' own order for two months | Refit on dispositions |
| Score PSI, monthly against development | above 0.25 | Revalidate |
| Alert mix by rule | any rule's share moves by half | Review the rule thresholds before the model |
Effective challenge
From DECISIONS.md:
| Date | Challenge | Response | What changed |
|---|---|---|---|
| 2026-09-27 | On months it never saw, the triage model put every alert that matched an injected typology ahead of every legitimate lookalike: an AUROC of one. | The generator writes its lookalikes and its criminals from separate recipes. A cash intensive business deposits cash every week at one branch; a structurer is new to cash and spreads deposits across several branches. Removing the cash history and the branch count still leaves the alert size, its entry count and the account's earlier alerts, and the separation survives, so it is not one leaked column. It is two synthetic populations that do not overlap. A real program's labels are investigator dispositions, which are noisier and arrive late. | The report publishes the ranking without the two strongest features and the rules' own order as benchmarks, states that the measured quality is a property of the generator, and makes investigator dispositions the monitoring measure the model would be refit on before any real use. |
| 2026-09-27 | The mule, rapid movement and round tripping rules fired no false alerts on this data, so a model that only learned which rule fired would already rank most of the queue correctly. | Pooled over all four rules, AUROC mostly rewards putting those three rules' alerts first, which each rule's own precision on the development months already does. The only alerts where order is a real question are structuring alerts, where the rule's precision is under half. | A gate now measures the model within structuring alerts alone, where the rules' order carries no information, and the benchmark table includes ranking by each rule's development precision. |
Promotion gates
| Gate | Threshold | Measured | Result |
|---|---|---|---|
| Test AUROC at or above the floor | at least 0.8 | 1.0000 | pass |
| Test average precision above the rules' own order | at least 0 | 0.1879 | pass |
| Test AUROC within structuring alerts, where the rules' order says nothing | at least 0.8 | 1.0000 | pass |
| Point in time recompute of sampled alerts | at least 200 | 200.0000 | pass |
| Expected calibration error on test | at most 0.1 | 0.0155 | pass |
| Test AUROC without the cash history and branch count features | at least 0.75 | 0.9968 | pass |
Validation conclusion
The model orders the queue the way its purpose asks and passes every gate on months it never saw. Its measured quality is a property of the generator, whose lookalikes and typologies do not overlap, so the approval is for ordering this demonstration's queue; on a real portfolio it would be refit and re-measured on investigator dispositions first.
approved. Conditions:
| Condition |
|---|
| None |
The integration contract
| Line | What this model does |
|---|---|
| trained | On alerts that closed before September 2025, labelled by the generator's injected truth. |
| timed | History features stop at the alert window's start; a leakage test recomputes 200 sampled alerts from raw entries. |
| calibrated | Reliability on the test months is published; the queue uses the score as a rank. |
| useful | Beats the base rate and the order the rules alone imply, including within structuring alerts, where the rules cannot rank at all. |
| fair | No demographic enters; the features are the alert's facts and the account's own history. |
| gated | Every gate in this report, including an ablation without the two features the generator made most different. |
| served | Scored in the pipeline; the ranked queue is published with the crime page. |
| integrated | Sorts the case queue: an investigator opens alerts in the model's order within each month. |
| monitored | Queue capture against investigator dispositions once they exist; monthly score PSI; alert mix by rule. |
| documented | This validation report and docs/crime.md. |
| bounded | Orders alerts; never closes one, never files anything, and never replaces an investigator's disposition. |
| validated | This report, rendered from the manifest before the queue uses the order. |
| explainable | Each alert shows the rule's statement in plain words and the features that moved its score most. |