Model risk · validation report · fraud_card

Validation report: Card transaction fraud (the desk's model)

Harborline Bank is fictional. This is a demonstration on generated data and public datasets; no real customer, account or transaction of any real institution appears here.

Nothing here is a credit decision, a fraud determination, a suspicious activity finding or investment advice.

Model fraud_card
Purpose Score every card authorisation at the moment it happens and decide approve, review or decline at a cost based operating point
Tier 1: Material: informs a credit, fraud or regulatory decision directly
Role sole
Status in use
Owner Head of analytics
Prediction time The instant of the authorisation; features use only the card's and the merchant's history strictly before it
Outcome Fraud as confirmed after the 60 day chargeback window
Last validated 2026-08-31
Conclusion approved

Purpose and scope

The desk's model scores every card authorisation the instant it happens. Above the operating point the transaction is declined; below it, a transaction that trips one of the desk's rules goes to review; everything else is approved. Alerts at or above the stricter queue point go to the analyst queue first. The model informs the decision in real time, so it is tier one.

Out of scope: application fraud (that is fraud_account), account takeover, and any decision about a customer rather than a transaction.

Conceptual soundness

Fraud is a stream, so the model is built around the moment of the transaction. Every feature uses only what the bank knew strictly before it: the card's transaction count and spend over the last ten minutes, hour, day and week; the time since its last purchase; whether this is the card's first purchase in the category; how the amount compares with the card's own history; the distance from the cardholder's home and from the previous merchant; and the merchant's fraud rate among labels that had arrived in the trailing ninety days. Rolling windows are closed on the left, so a transaction never counts itself.

A gradient boosted model suits a problem with sharp interactions between amount, category and time. A logistic model on the same features is the benchmark, and the base rate is the floor. The operating point is chosen by cost, not accuracy: a false decline costs $12.00, the interchange forgone and the customer's call; a missed fraud costs its amount plus $35.00 of handling. The threshold minimising expected cost on the validation months is 0.0503. Accuracy would have chosen 0.6376, a far stricter cut, because accuracy counts a missed fraud and a declined coffee as the same mistake. On the test months that choice costs $89,153 against $24,927 at the cost based point.

The queue point, 0.2759, is the lowest threshold whose validation precision reached 90%, and never below the operating point. A queued alert ends in a card block and a call to the customer, which is heavier than declining one transaction, so the queue holds only what the model is surest of, and an analyst's hour goes to the alerts most likely to need it.

Data and assumptions

Card transactions generated by Sparkov (MIT) with a seed, identities removed at ingest. Trained on September 2025 to March 2026, whose labels had all arrived before the thresholds were chosen; the operating and queue points were chosen on April and May 2026; everything below is reported on June to August 2026, 303,401 transactions with 2,816 frauds, which the thresholds never saw.

Labels arrive late. A transaction's fraud label is known only after the 60 day chargeback window. On the as of date, 200,569 of the test months' transactions (66.1%) were still inside it. The test figures in this report use the labels as they eventually arrive, because the generator knows them; the desk's own figures never do.

A leakage test recomputed 200 sampled rows' features from raw history, strictly before each transaction, and matched every one. Cardholder gender and age are in the data for the fairness audit and are on the proxy list; the cardholder's coordinates and city population are proxies too. None is a feature.

Development evidence

Early stopped on the validation months after 234 trees. Development AUROC 1.000, validation 0.999, and average precision 0.996 and 0.981.

Outcomes analysis on held out data

On the test months the model reaches an AUROC of 1.000 and an average precision of 0.978.

At the operating point it declines 3,676 transactions, with precision 74.8% and recall 97.6%, at a false positive rate of 0.308%. The expected cost on the test months is $24,927 against $1,573,327 for approving everything: a saving of $1,548,400 (98.4%), measured on months the threshold never saw.

point threshold declined_or_queued cost
Cost minimum (operating point) 0.0503 3,676 $24,927
Queue point 0.2759 2,970 not applicable
Accuracy maximum 0.6376 2,659 $89,153
Approve everything not applicable 0 $1,573,327

At the queue point the queue holds 2,970 alerts at precision 89.7%, catching 94.6% of fraud.

Late labels. On the as of date the desk could compute precision only over decisions older than the chargeback window: 73.4%, with recall 98.1%. A dashboard that counted every unlabelled decline as a false positive would have reported precision of 27.0% for the same period. That second number is wrong by construction, and it is the number most dashboards show.

The rules, on the test months:

rule fired precision recall
First purchase over $500 in a category this card has never used 88 100.0% 3.1%
Three or more purchases on this card within 10 minutes 272 12.1% 1.2%
Merchant more than 220 km from this card's previous merchant within the hour 521 1.7% 0.3%

By cardholder gender and age band at the operating point:

Group N Event rate Mean score Calibration gap Auroc Flag rate Fpr Recall
F 155,702 0.95% 0.96% 0.0% 0.999 1.27% 0.36% 96.4%
M 147,699 0.90% 0.94% 0.0% 1.000 1.14% 0.26% 98.9%
25 to 39 93,649 0.75% 0.75% 0.0% 0.999 1.00% 0.29% 95.6%
40 to 59 116,789 0.88% 0.90% 0.0% 0.999 1.19% 0.33% 97.1%
60 and over 63,009 1.27% 1.33% +0.1% 1.000 1.59% 0.34% 99.4%
under 25 29,954 0.96% 0.98% 0.0% 1.000 1.16% 0.20% 99.7%

The lowest false positive rate across age bands is 0.60 of the highest.

Calibration on the test sample, ten equal count bins:

Bin Accounts Predicted Observed
1 30,341 0.00% 0.00%
2 30,340 0.00% 0.00%
3 30,340 0.00% 0.00%
4 30,340 0.00% 0.00%
5 30,340 0.00% 0.00%
6 30,340 0.00% 0.00%
7 30,340 0.00% 0.00%
8 30,340 0.00% 0.00%
9 30,340 0.01% 0.00%
10 30,340 9.49% 9.28%

Source: replayed, Sparkov card transactions, September 2025 to August 2026, as of 2026-08-31.

Benchmarking

Against the base rate and a logistic model on the same features, on the same test months.

Model AUROC Gini Brier ECE
This model 1.000 0.999 0.0010 0.0002
Base rate 0.500 0.000 0.0092 0.0031
Logistic on the same features 0.933 0.866 0.0053 0.0015

Source: replayed, Sparkov card transactions, September 2025 to August 2026, as of 2026-08-31.

Limitations

  • Sparkov's fraud is easy to separate. Fraudulent amounts, hours and categories come from distributions that barely overlap with genuine ones, so the discrimination figures describe the generator more than any real fraud problem. What carries over is the operating machinery: the cost based threshold, the queue point, late labels and labelled only precision.
  • Chargebacks are simulated by a fixed window. Real labels arrive over a distribution of delays, and some never arrive.
  • The false decline cost is a stated assumption; docs/fraud-desk.md defends it and the cost curve shows how the operating point moves if it is wrong.
  • The distant merchant rule has little to find in Sparkov's geography.

Ongoing monitoring plan

On the desk and weekly:

Measure Trigger Action
Precision over labelled decisions, weekly below 90% at the queue point for two weeks Review the threshold on the latest labelled month
Unlabelled decisions inside the chargeback window reported beside every precision figure None: it is context, not a trigger
Score PSI, weekly against validation above 0.25 Revalidate
Rule precision on labelled months any rule below twice the base rate Retire or retune the rule

Effective challenge

From DECISIONS.md:

Date Challenge Response What changed
2026-09-26 The merchant features were a cumulative count of a merchant's labelled transactions and a fraud rate over every label since the data began. Both grow or settle with the calendar, so the same merchant looks different in March than in August for no reason a fraud analyst would accept. A feature whose meaning drifts with elapsed time teaches a model the calendar. On the desk it would also make a new merchant look risky forever and an old one safe forever. The merchant fraud rate now uses only labels that arrived in the trailing 90 days, smoothed toward the rate over the same window, and the raw volume count was removed from the model.
2026-09-26 At 150 km the rule fired on roughly one transaction in thirty, at about twice the base fraud rate. Every one of those is a review an analyst has to clear. Sparkov places merchants at random within a band around the cardholder, so consecutive merchants are often far apart for genuine reasons. A rule that doubles the base rate while flooding the queue costs more analyst time than it saves. The threshold moved to 220 km, which cut the rule's volume by more than an order of magnitude. Its precision is published in the validation report, and the monitoring plan retires any rule that falls below twice the base rate.
2026-09-26 Once the stream was corrected, the model separates fraud from genuine transactions almost perfectly on the test months. That is a property of the generator, not an achievement. Sparkov draws fraudulent amounts, hours and categories from distributions that barely overlap with genuine ones; real card fraud imitates genuine spending. What carries over to a real desk is the machinery around the score: the cost based operating point, the queue point, labels that arrive late, and precision that only counts what has been labelled. The validation report and the desk page lead with cost and late labels, state this limitation beside the discrimination figures, and the AUROC gate is not the reason the model is approved.

Promotion gates

Gate Threshold Measured Result
Test AUROC at or above the floor at least 0.9 0.9996 pass
Test average precision above the logistic benchmark at least 0 0.4250 pass
Test cost at the operating point below approving everything (share saved) at least 0.3 0.9842 pass
Lowest over highest false positive rate across age bands at the operating point at least 0.5 0.6020 pass
Point in time recompute of sampled rows at least 200 200.0000 pass
Expected calibration error on test at most 0.01 0.0002 pass

Validation conclusion

The model saves most of the cost of fraud on months it never saw, its operating point is chosen the way a fraud operations team would choose it, and its live figures respect late labels. Its false positive rates by age band are reported with the condition below if they are not within the stated band. Approved for the desk, with the discrimination figures read as a property of the generator.

approved. Conditions:

Condition
None

The integration contract

Line What this model does
trained On September 2025 to March 2026, whose labels had all arrived by the end of May.
timed Every feature is computed from strictly earlier transactions; a leakage test recomputes 200 sampled rows from raw history.
calibrated Reliability on the test months is published; the cost curve uses scores as ranks, not as probabilities.
useful Beats the base rate and a logistic model on the same features, and saves most of the cost of approving everything.
fair False positive rate and recall by cardholder gender and age band at the operating point.
gated Every gate in its validation report, including the cost saving measured on data the threshold never saw.
served POST /v1/score/fraud_card and the desk's replay worker in the API; recorded replay on the site.
integrated Decides approve, review or decline on the desk; alerts at the queue point go to the analyst queue first.
monitored Precision and recall over labelled transactions only, with the unlabelled count beside them; weekly score PSI.
documented This validation report, docs/fraud-desk.md and the rule file.
bounded Scores card authorisations only; the rules can send a transaction to review but never decline it alone.
validated This report, rendered from the manifest before the desk uses the model.
explainable Every review or decline shows the rules that fired and the two features that moved the score most, in plain words.