Function 10 · model risk management

Which models make decisions here, and has each one earned it?

The inventory holds 8 models, each with a validation report written in the order SR 11-7 asks for: purpose, data, conceptual soundness, outcomes analysis, benchmarks, limitations, monitoring and a conclusion. The conclusions: approved 4, approved with conditions 2, not approved 2. The challenge log holds 17 challenges to the models, each with what it changed.

Models in the inventory
8
Worst application PSI
0.005
In 2026Q3
Worst weekly fraud score PSI
0.031
Weeks before fraud AUROC could be measured
8
First labelled week 2026-07-27

The inventory

aml_triage
Alert triage for transaction monitoring
Rank the transaction monitoring rules' alerts so an investigator opens the likeliest real typology first
Tier
2
Role
sole
Status
in use
Test AUROC
1.000
Last validated
2026-08-31
Conclusion
✓ approved
Read the validation report
attrition
Deposit customer attrition, twelve month horizon
Rank customers by the balance the bank stands to lose if they leave in the next twelve months, so retention calls go where they are worth most
Tier
2
Role
sole
Status
in use
Test AUROC
0.583
Last validated
2026-08-31
Conclusion
✓ approved
Read the validation report
cure
Cure within ninety days for delinquent loans
Order the collections queue by the balance likely to be lost without a call: balance times the probability the loan does not cure
Tier
2
Role
sole
Status
in use
Test AUROC
0.787
Last validated
2026-08-31
Conclusion
✓ approved
Read the validation report
fraud_account
Application fraud on the BAF suite
Score a bank account application for fraud when it is submitted, so that the riskiest ask for documents before the account opens
Tier
1
Role
sole
Status
in use
Test AUROC
0.880
Last validated
2026-08-31
Conclusion
⚠ approved with conditions
  • Remediate before the next validation: Predictive equality between the age groups at the suite's threshold (measured 0.4558)
Read the validation report
fraud_card
Card transaction fraud (the desk's model)
Score every card authorisation at the moment it happens and decide approve, review or decline at a cost based operating point
Tier
1
Role
sole
Status
in use
Test AUROC
1.000
Last validated
2026-08-31
Conclusion
✓ approved
Read the validation report
pd_challenger
Probability of default challenger (monotone GBM)
Test whether a gradient boosted model beats the scorecard enough to justify its opacity
Tier
1
Role
challenger
Status
validated
Test AUROC
0.667
Last validated
2026-08-31
Conclusion
✕ not approved
  • Test AUROC at least 0.01 above the champion (measured 0.00535)
Read the validation report
pd_scorecard
Probability of default scorecard (champion)
Rank and price installment loan applications by twelve month default risk and state the reasons for a decline
Tier
1
Role
champion
Status
in use
Test AUROC
0.662
Last validated
2026-08-31
Conclusion
⚠ approved with conditions
  • Remediate before the next validation: Expected calibration error on test (measured 0.02859)
Read the validation report
relief
Monetary relief on consumer complaints
Route complaints likely to end in a refund to the team that can fix the cause, and rank issues by the relief they are likely to cost
Tier
2
Role
sole
Status
development
Test AUROC
0.869
Last validated
2026-08-31
Conclusion
✕ not approved
  • Test AUROC above the metadata only model: the narrative has to add something (measured 0.01401)
Read the validation report

Score stability by application quarter, against the development quarters

0.000.050.100.150.200.252023Q32024Q12024Q32025Q12025Q32026Q12026Q3Application quarterBreachWarningPSI

Population stability index

The worst quarter's population stability index is 0.005 (2026Q3), against a warning line at 0.10 and a breach at 0.25. Stability says the applicants look alike; it says nothing about whether they default alike, which is why calibration is monitored separately as outcomes mature.

Source: generated, every installment application, approved or declined, seed 20260831, as of 2026-08-31.

The fraud model's monitoring, week by week

0.0000.0050.0100.0150.0200.0250.0300.0352026-06-012026-06-152026-06-292026-07-132026-07-272026-08-102026-08-24Weekfirst labelled weekScore PSI

Score PSI against validation

Score stability is measurable the day a transaction is scored; discrimination is not. For the first 8 weeks of the test months no test transaction had cleared the 60 day chargeback window, so the model's AUROC on the new months could not be computed at all. That gap is the monitoring lag, and the plan says so rather than filling it.

Source: replayed, fraud_card scored stream, April to August 2026, as of 2026-08-31.

The challenge log

Every time a result was questioned, what was found and what changed because of it.

DateModelsChallengeWhat changed
2026-09-26pd_scorecardOn the validation months the scorecard's mean predicted default rate sat well under the observed rate, while its rank ordering held up.Calibration became a non critical gate, so the report concludes approved with a stated condition rather than approved, and expected loss in function 2 applies an observed to expected overlay by vintage instead of the raw scorecard probability.
2026-09-26pd_scorecardThe fitted scorecard has no delinquency characteristic. Delinquencies in the last 24 months fell under the information value floor and were dropped.The published information value table now lists every candidate with the reason it was kept or dropped, so a reviewer sees the decision instead of an absence, and the challenger, which sees every candidate, confirms the characteristic adds little.
2026-09-26pd_challengerThe first fit's development AUROC was far above its validation AUROC.Seven leaves and a minimum of 150 loans per leaf. The development to validation gap narrowed and the validation and test AUROC held.
2026-09-26pd_challengerThe challenger's test AUROC is above the champion's.Promotion requires a test AUROC at least 0.01 above the champion's, stated as a gate in the validation report, together with the stability and age band gates.
2026-09-26fraud_cardThe merchant features were a cumulative count of a merchant's labelled transactions and a fraud rate over every label since the data began. Both grow or settle with the calendar, so the same merchant looks different in March than in August for no reason a fraud analyst would accept.The merchant fraud rate now uses only labels that arrived in the trailing 90 days, smoothed toward the rate over the same window, and the raw volume count was removed from the model.
2026-09-26fraud_cardAt 150 km the rule fired on roughly one transaction in thirty, at about twice the base fraud rate. Every one of those is a review an analyst has to clear.The threshold moved to 220 km, which cut the rule's volume by more than an order of magnitude. Its precision is published in the validation report, and the monitoring plan retires any rule that falls below twice the base rate.
2026-09-26fraud_cardOnce the stream was corrected, the model separates fraud from genuine transactions almost perfectly on the test months.The validation report and the desk page lead with cost and late labels, state this limitation beside the discrimination figures, and the AUROC gate is not the reason the model is approved.
2026-09-27aml_triageOn months it never saw, the triage model put every alert that matched an injected typology ahead of every legitimate lookalike: an AUROC of one.The report publishes the ranking without the two strongest features and the rules' own order as benchmarks, states that the measured quality is a property of the generator, and makes investigator dispositions the monitoring measure the model would be refit on before any real use.
2026-09-27aml_triageThe mule, rapid movement and round tripping rules fired no false alerts on this data, so a model that only learned which rule fired would already rank most of the queue correctly.A gate now measures the model within structuring alerts alone, where the rules' order carries no information, and the benchmark table includes ranking by each rule's development precision.
2026-09-27fraud_accountcustomer_age is on the proxy list and never a feature, yet at the suite's threshold older applicants who are not fraudsters are flagged far more often than younger ones. If the model cannot see age, where does the gap come from?The date of birth count is on the proxy list now, with its reason, so no model may use it. The report publishes a search for a less discriminatory alternative: the next strongest age carriers removed one at a time, with the recall each step costs and the parity it buys, on both variants. The model's use is bounded to asking for documents before an account opens, never to declining one, because a false flag then costs an applicant a delay rather than an account.
2026-09-27fraud_accountRecall at a five percent false positive rate sets the threshold on the same months it reports. A bank has to fix its threshold before the applications arrive, and the false positive rate it then gets is whatever the new months give it.The report and the page publish both: the suite's measure, and the recall, false positive rate and predictive equality on months 6 and 7 at the threshold fixed on month 5. The recall gate uses the fixed threshold, and a gate watches how far the realised false positive rate drifts from its target.
2026-09-27attritionThe first split trained on every month end up to December 2024 and tested on the summer of 2025. A December 2024 label says whether the customer left by December 2025, which the bank could not have known when it scored the summer of 2025.Training uses month ends September 2023 to February 2024, validation March to May 2024, and test June to August 2025, so every training and validation outcome was known before the first test month end. The gap is stated in the report beside the split.
2026-09-27attritionSorted by probability alone, the top of the retention list fills with new customers holding a few hundred dollars. A retention call costs the same whoever answers.The list is ranked by expected balance at risk. The report publishes, for the top tenth of the test months, the share of closers and the share of their balances found by each ordering, so the trade between the two is visible.
2026-09-27cureA loan thirty days past due cures far more often than one ninety days past due. A model that only learned the bucket would beat the base rate comfortably and add nothing a collections supervisor does not already do by working the thirty day list first.The bucket rule is a benchmark in the report, and a critical gate requires the model's test AUROC to beat the rule's by two hundredths. Cure rates by bucket are published beside the model's so the rule's strength is visible.
2026-09-27cureA loan that is past due in June and July 2025 appears in the last training month and the first validation month, with nearly the same features and the same outcome. Early stopping and the choice between the two candidates were being made on loans the fit had already seen.One loan in five is held out of the fit by a hash of its identifier, and early stopping and the choice of champion use only held out loans in the validation months. The test months begin three months after validation ends, so every outcome used to fit or choose the model was known before the first test month end, and a check in the trainer refuses to run otherwise.
2026-09-27reliefComplaints received in late 2024 ended in monetary relief about one time in seventy; by mid 2026 it was closer to one in fourteen. A model fitted on the early months will score the later ones too low, and a probability threshold chosen on 2025 would flag almost nothing in 2026.The operating point flags the top twentieth of scores, set on the validation months immediately before the test months. Reliability on the test months is published with the drift stated beside it, the calibration gate is non critical, and the monitoring plan resets the operating point when the monthly base rate moves by half.
2026-09-27reliefConsumers write about their age, their disability, their military service and their family. A text model can learn that an elderly widow is more likely to get a refund, which is a protected basis entering a routing decision by the back door. The Bureau's own tags mark older Americans and servicemembers.`rules/proxies.yaml` now lists words a text model may not learn from, and the relief model drops them, and any phrase containing them, from its vocabulary. The Bureau's tags are never features; they are used only to audit flag rates by group, and a gate watches the ratio.

Characteristic stability, scorecard

CharacteristicStability index, last six months against development
purpose0.003
oldest_trade_months0.003
revolving_utilization0.002
bureau_score0.000
dti0.000
employment_months0.000
loan_amount0.000
inquiries_6m0.000

Source: generated, every installment application, approved or declined, seed 20260831, as of 2026-08-31.