Model risk · validation report · fraud_account
Validation report: Application fraud on the BAF suite
Harborline Bank is fictional. This is a demonstration on generated data and public datasets; no real customer, account or transaction of any real institution appears here.
Nothing here is a credit decision, a fraud determination, a suspicious activity finding or investment advice.
| Model | fraud_account |
| Purpose | Score a bank account application for fraud when it is submitted, so that the riskiest ask for documents before the account opens |
| Tier | 1: Informs a decision a person makes, with a method whose behaviour is hard to inspect |
| Role | sole |
| Status | in use |
| Owner | Head of analytics |
| Prediction time | When the application is submitted; the suite's features describe the application and the activity before it |
| Outcome | The application was fraudulent, as the suite labels it |
| Last validated | 2026-08-31 |
| Conclusion | approved with conditions |
Purpose and scope
Score a bank account application for fraud when it is submitted, so that the riskiest applications ask for documents before the account opens. The model informs a check a person completes; it never declines an application. It is tier one because a gradient boosted model's behaviour is hard to inspect, and because its errors fall unevenly by age.
In scope: the Bank Account Fraud suite (Jesus et al., NeurIPS 2022), a set of synthetic
account applications generated from a real fraud dataset and published with the
protected attributes kept so fairness can be measured. This report reproduces the
suite's evaluation protocol on two of its variants. Out of scope: card fraud, which is
fraud_card, and any decision to refuse an account.
Conceptual soundness
Application fraud is rare and the features are a mix of counts, durations, flags and categories with missing values marked by negative numbers, which gradient boosted trees handle without imputation. The suite's own measures are used as the suite defines them: recall at a 5% false positive rate, and predictive equality, the lower of the two age groups' false positive rates over the higher at that threshold, with the groups split at 50. Parity is one.
Age is measured and never used. customer_age is a protected attribute on the proxy
list, as is name_email_similarity, which is computed from the applicant's name, and
the feature list is checked against that list before training.
Data and assumptions
A 300,000 application sample of the suite's base variant, and a sample of
the same size of variant II beside it. The suite trains on months 0 to 5
and tests on months 6 and 7. Here the model trains on months 0 to 4,
month 5 chooses the tree count and the threshold a bank could deploy,
and the test is the suite's, so every figure below comes from months the model and its
threshold never saw. The trade is one month of training data, recorded in DECISIONS.md.
The suite computes its features at application time and ships no raw history, so there
is nothing to recompute a feature from. The leakage check is a screen instead: a feature
known only after the outcome would predict it almost alone, and the strongest single
feature here, housing_status, reaches a test AUROC of 0.724.
Development evidence
Development AUROC 0.950, validation 0.882, early stopped at 143 trees.
Outcomes analysis on held out data
On the test months the model reaches an AUROC of 0.880. At the suite's threshold it finds 49.0% of fraudulent applications. Applicants aged 50 and over who are not fraudsters are flagged at 9.19%, against 4.19% for the rest, 2.2 times as often: a predictive equality of 0.46.
With the threshold fixed on month 5 instead, as a bank would have to fix it, the test months give a recall of 53.8% at a false positive rate of 5.94% and a predictive equality of 0.47.
The two variants side by side. Variant II's age groups are about the same size (50.5% of its applicants are 50 or over, against 18.3% in the base), and its older applicants' fraud rate is 4.7 times the younger group's, against 2.9 times in the base:
| Measure | Base | Variant II |
|---|---|---|
| Applications in the sample | 300,000 | 300,000 |
| Share aged fifty and over | 18.3% | 50.5% |
| Fraud rate, aged fifty and over | 2.34% | 1.82% |
| Fraud rate, under fifty | 0.81% | 0.38% |
| Test AUROC | 0.880 | 0.882 |
| Recall at a five percent false positive rate (the suite's measure) | 49.0% | 49.3% |
| False positive rate, aged fifty and over | 9.19% | 6.96% |
| False positive rate, under fifty | 4.19% | 3.04% |
| Predictive equality (lower rate over higher) | 0.46 | 0.44 |
| Logistic benchmark: recall at the same rate | 47.2% | 47.7% |
| Threshold fixed on month five: recall on the test months | 53.8% | 52.9% |
| Threshold fixed on month five: false positive rate on the test months | 5.94% | 6.07% |
| Threshold fixed on month five: predictive equality | 0.47 | 0.44 |
| AUROC of the features at predicting age fifty and over | 0.782 | 0.788 |
Source: real:baf, Bank account applications in the BAF suite's base variant, a 300,000 sample; variant II beside it, as of 2026-08-31.
By age band, at the suite's threshold on the base variant:
| Group | N | Event rate | Mean score | Calibration gap | Auroc | Flag rate | Fpr | Recall |
|---|---|---|---|---|---|---|---|---|
| under 30 | 15,411 | 0.66% | 0.55% | -0.1% | 0.883 | 2.42% | 2.19% | 36.6% |
| 30 to 39 | 18,963 | 1.19% | 0.82% | -0.4% | 0.858 | 4.38% | 3.91% | 43.8% |
| 40 to 49 | 16,755 | 1.42% | 1.13% | -0.3% | 0.874 | 6.97% | 6.36% | 49.2% |
| 50 to 59 | 7,622 | 2.45% | 1.56% | -0.9% | 0.876 | 10.43% | 9.21% | 58.8% |
| 60 and over | 2,434 | 4.15% | 1.66% | -2.5% | 0.860 | 11.01% | 9.13% | 54.5% |
Source: real:baf, Bank account applications in the BAF suite's base variant, a 300,000 sample; variant II beside it, as of 2026-08-31.
Calibration on the test sample, ten equal count bins:
| Bin | Accounts | Predicted | Observed |
|---|---|---|---|
| 1 | 6,119 | 0.06% | 0.02% |
| 2 | 6,119 | 0.09% | 0.10% |
| 3 | 6,119 | 0.13% | 0.15% |
| 4 | 6,119 | 0.17% | 0.21% |
| 5 | 6,119 | 0.22% | 0.33% |
| 6 | 6,118 | 0.30% | 0.44% |
| 7 | 6,118 | 0.41% | 0.47% |
| 8 | 6,118 | 0.63% | 1.03% |
| 9 | 6,118 | 1.18% | 2.16% |
| 10 | 6,118 | 6.44% | 9.04% |
Source: real:baf, Bank account applications in the BAF suite's base variant, a 300,000 sample; variant II beside it, as of 2026-08-31.
Benchmarking
Against the base rate and a logistic regression on the same features, with a flag for each missing marker and the categories one hot encoded. At the suite's measure the logistic model's recall is 47.2%.
| Model | AUROC | Gini | Brier | ECE |
|---|---|---|---|---|
| This model | 0.880 | 0.760 | 0.0126 | 0.0044 |
| Base rate | 0.500 | 0.000 | 0.0138 | 0.0042 |
| Logistic on the same features | 0.875 | 0.750 | 0.0129 | 0.0069 |
Source: real:baf, Bank account applications in the BAF suite's base variant, a 300,000 sample; variant II beside it, as of 2026-08-31.
Limitations
The age gap does not close by leaving age out. The model's features predict whether an applicant is 50 or over with an AUROC of 0.782; the strongest carriers after the date of birth count left the feature list are employment_status, housing_status, current_address_months_count. Removing them one at a time is the search for a less discriminatory alternative:
| variant | removed | auroc | recall | fpr_older | fpr_younger | ratio |
|---|---|---|---|---|---|---|
| Base | Nothing (the model) | 0.880 | 49.0% | 9.19% | 4.19% | 0.46 |
| Base | employment_status | 0.877 | 48.3% | 8.93% | 4.24% | 0.48 |
| Base | employment_status, housing_status | 0.859 | 44.8% | 7.89% | 4.44% | 0.56 |
| Base | employment_status, housing_status, current_address_months_count | 0.854 | 44.2% | 7.80% | 4.46% | 0.57 |
| Variant II | Nothing (the model) | 0.882 | 49.3% | 6.96% | 3.04% | 0.44 |
| Variant II | employment_status | 0.879 | 49.5% | 6.71% | 3.29% | 0.49 |
| Variant II | employment_status, housing_status | 0.860 | 44.3% | 6.62% | 3.38% | 0.51 |
| Variant II | employment_status, housing_status, current_address_months_count | 0.855 | 43.4% | 6.45% | 3.55% | 0.55 |
Source: real:baf, Bank account applications in the BAF suite's base variant, a 300,000 sample; variant II beside it, as of 2026-08-31.
Parity is bought with recall: removing all three moves the base variant's predictive
equality to 0.57 at a recall of 44.2%. The
model is kept, and bounded to asking for documents, because a false flag then costs an
applicant a delay, not an account. The choice is recorded in DECISIONS.md and the
search is rerun at every revalidation.
- The suite is synthetic. Its features are generated to match a real dataset's joint distribution, not produced by a real application flow, so a real bank's separability and drift will differ.
- The suite's labels are final; a real application fraud label arrives weeks later, when the account is used.
Ongoing monitoring plan
Monthly, once labels arrive:
| Measure | Trigger | Action |
|---|---|---|
| False positive rate at the threshold, monthly, by age group | overall above 7.5% or the ratio below its validated value | Reset the threshold on the latest labelled month |
| Recall on labelled applications, monthly | below the validated recall for two months | Revalidate |
| Score PSI, monthly against month five | above 0.25 | Revalidate |
Effective challenge
From DECISIONS.md:
| Date | Challenge | Response | What changed |
|---|---|---|---|
| 2026-09-27 | customer_age is on the proxy list and never a feature, yet at the suite's threshold older applicants who are not fraudsters are flagged far more often than younger ones. If the model cannot see age, where does the gap come from? | From the other features. A model trained to predict fifty and over from the fraud model's own features separates the two groups well, and the strongest single carrier was date_of_birth_distinct_emails_4w: a count taken by grouping applications on the applicant's date of birth, so its value depends on how common the applicant's birth year is among applicants. Employment status, housing status and time at the current address come next, and each is also a legitimate fraud signal. | The date of birth count is on the proxy list now, with its reason, so no model may use it. The report publishes a search for a less discriminatory alternative: the next strongest age carriers removed one at a time, with the recall each step costs and the parity it buys, on both variants. The model's use is bounded to asking for documents before an account opens, never to declining one, because a false flag then costs an applicant a delay rather than an account. |
| 2026-09-27 | Recall at a five percent false positive rate sets the threshold on the same months it reports. A bank has to fix its threshold before the applications arrive, and the false positive rate it then gets is whatever the new months give it. | The suite's measure is right for comparing models with each other and with the published results, and wrong as a statement of what the bank would see. | The report and the page publish both: the suite's measure, and the recall, false positive rate and predictive equality on months 6 and 7 at the threshold fixed on month 5. The recall gate uses the fixed threshold, and a gate watches how far the realised false positive rate drifts from its target. |
Promotion gates
| Gate | Threshold | Measured | Result |
|---|---|---|---|
| Test AUROC at or above the floor | at least 0.8 | 0.8801 | pass |
| Test recall at the threshold fixed on month five | at least 0.4 | 0.5381 | pass |
| Recall at a five percent false positive rate above the logistic benchmark | at least 0 | 0.0176 | pass |
| Test false positive rate at the threshold fixed on month five | at most 0.075 | 0.0594 | pass |
| Strongest single feature's test AUROC (a feature known after the outcome would stand alone) | at most 0.8 | 0.7235 | pass |
| Predictive equality between the age groups at the suite's threshold | at least 0.8 | 0.4558 | FAIL |
| Expected calibration error on test | at most 0.01 | 0.0044 | pass |
Validation conclusion
The model does the job its purpose states and passes every critical gate on months it never saw, at a threshold fixed before them. It does not reach predictive equality between the age groups on either variant, which is why it is approved with that condition, bounded to routing applications to document checks, and published with the alternatives that would trade recall for parity.
approved with conditions. Conditions:
| Condition |
|---|
| Remediate before the next validation: Predictive equality between the age groups at the suite's threshold (measured 0.4558) |
The integration contract
| Line | What this model does |
|---|---|
| trained | On months 0 to 4 of the base variant; month 5 chooses the tree count and the deployable threshold. |
| timed | The suite computes its features at application time and ships no raw history to recompute them from, so the split is temporal and a single feature screen looks for anything known only after the outcome. |
| calibrated | Reliability on the test months is published; the threshold is set by false positive rate, so the score is used as a rank. |
| useful | Beats the base rate and a logistic model on the same features at the suite's own measure, recall at a five percent false positive rate. |
| fair | Predictive equality between applicants aged fifty and over and the rest, on both variants, with a search for a less discriminatory alternative published beside it. |
| gated | Every gate in this report, including the recall at a threshold fixed before the test months. |
| served | Scored in the pipeline; not served by the API, because no live application stream exists here. |
| integrated | An application at or above the threshold asks for documents before the account opens. The model never declines one. |
| monitored | Monthly false positive rate against the threshold, by age group; recall on labelled applications; score PSI. |
| documented | This validation report and the desk page's fairness audit. |
| bounded | Routes applications to document checks only; never declines, never prices, and never sees age. |
| validated | This report, rendered from the manifest before the threshold is used. |
| explainable | A routed application shows the features that moved its score most, in the suite's own terms. |