Function 10 · model risk management
Which models make decisions here, and has each one earned it?
The inventory holds 8 models, each with a validation report written in the order SR 11-7 asks for: purpose, data, conceptual soundness, outcomes analysis, benchmarks, limitations, monitoring and a conclusion. The conclusions: approved 4, approved with conditions 2, not approved 2. The challenge log holds 17 challenges to the models, each with what it changed.
The inventory
- Tier
- 2
- Role
- sole
- Status
- in use
- Test AUROC
- 1.000
- Last validated
- 2026-08-31
- Conclusion
- ✓ approved
- Tier
- 2
- Role
- sole
- Status
- in use
- Test AUROC
- 0.583
- Last validated
- 2026-08-31
- Conclusion
- ✓ approved
- Tier
- 2
- Role
- sole
- Status
- in use
- Test AUROC
- 0.787
- Last validated
- 2026-08-31
- Conclusion
- ✓ approved
- Tier
- 1
- Role
- sole
- Status
- in use
- Test AUROC
- 0.880
- Last validated
- 2026-08-31
- Conclusion
- ⚠ approved with conditions
- Remediate before the next validation: Predictive equality between the age groups at the suite's threshold (measured 0.4558)
- Tier
- 1
- Role
- sole
- Status
- in use
- Test AUROC
- 1.000
- Last validated
- 2026-08-31
- Conclusion
- ✓ approved
- Tier
- 1
- Role
- challenger
- Status
- validated
- Test AUROC
- 0.667
- Last validated
- 2026-08-31
- Conclusion
- ✕ not approved
- Test AUROC at least 0.01 above the champion (measured 0.00535)
- Tier
- 1
- Role
- champion
- Status
- in use
- Test AUROC
- 0.662
- Last validated
- 2026-08-31
- Conclusion
- ⚠ approved with conditions
- Remediate before the next validation: Expected calibration error on test (measured 0.02859)
- Tier
- 2
- Role
- sole
- Status
- development
- Test AUROC
- 0.869
- Last validated
- 2026-08-31
- Conclusion
- ✕ not approved
- Test AUROC above the metadata only model: the narrative has to add something (measured 0.01401)
Score stability by application quarter, against the development quarters
Population stability index
Source: generated, every installment application, approved or declined, seed 20260831, as of 2026-08-31.
The fraud model's monitoring, week by week
Score PSI against validation
Source: replayed, fraud_card scored stream, April to August 2026, as of 2026-08-31.
The challenge log
Every time a result was questioned, what was found and what changed because of it.
| Date | Models | Challenge | What changed |
|---|---|---|---|
| 2026-09-26 | pd_scorecard | On the validation months the scorecard's mean predicted default rate sat well under the observed rate, while its rank ordering held up. | Calibration became a non critical gate, so the report concludes approved with a stated condition rather than approved, and expected loss in function 2 applies an observed to expected overlay by vintage instead of the raw scorecard probability. |
| 2026-09-26 | pd_scorecard | The fitted scorecard has no delinquency characteristic. Delinquencies in the last 24 months fell under the information value floor and were dropped. | The published information value table now lists every candidate with the reason it was kept or dropped, so a reviewer sees the decision instead of an absence, and the challenger, which sees every candidate, confirms the characteristic adds little. |
| 2026-09-26 | pd_challenger | The first fit's development AUROC was far above its validation AUROC. | Seven leaves and a minimum of 150 loans per leaf. The development to validation gap narrowed and the validation and test AUROC held. |
| 2026-09-26 | pd_challenger | The challenger's test AUROC is above the champion's. | Promotion requires a test AUROC at least 0.01 above the champion's, stated as a gate in the validation report, together with the stability and age band gates. |
| 2026-09-26 | fraud_card | The merchant features were a cumulative count of a merchant's labelled transactions and a fraud rate over every label since the data began. Both grow or settle with the calendar, so the same merchant looks different in March than in August for no reason a fraud analyst would accept. | The merchant fraud rate now uses only labels that arrived in the trailing 90 days, smoothed toward the rate over the same window, and the raw volume count was removed from the model. |
| 2026-09-26 | fraud_card | At 150 km the rule fired on roughly one transaction in thirty, at about twice the base fraud rate. Every one of those is a review an analyst has to clear. | The threshold moved to 220 km, which cut the rule's volume by more than an order of magnitude. Its precision is published in the validation report, and the monitoring plan retires any rule that falls below twice the base rate. |
| 2026-09-26 | fraud_card | Once the stream was corrected, the model separates fraud from genuine transactions almost perfectly on the test months. | The validation report and the desk page lead with cost and late labels, state this limitation beside the discrimination figures, and the AUROC gate is not the reason the model is approved. |
| 2026-09-27 | aml_triage | On months it never saw, the triage model put every alert that matched an injected typology ahead of every legitimate lookalike: an AUROC of one. | The report publishes the ranking without the two strongest features and the rules' own order as benchmarks, states that the measured quality is a property of the generator, and makes investigator dispositions the monitoring measure the model would be refit on before any real use. |
| 2026-09-27 | aml_triage | The mule, rapid movement and round tripping rules fired no false alerts on this data, so a model that only learned which rule fired would already rank most of the queue correctly. | A gate now measures the model within structuring alerts alone, where the rules' order carries no information, and the benchmark table includes ranking by each rule's development precision. |
| 2026-09-27 | fraud_account | customer_age is on the proxy list and never a feature, yet at the suite's threshold older applicants who are not fraudsters are flagged far more often than younger ones. If the model cannot see age, where does the gap come from? | The date of birth count is on the proxy list now, with its reason, so no model may use it. The report publishes a search for a less discriminatory alternative: the next strongest age carriers removed one at a time, with the recall each step costs and the parity it buys, on both variants. The model's use is bounded to asking for documents before an account opens, never to declining one, because a false flag then costs an applicant a delay rather than an account. |
| 2026-09-27 | fraud_account | Recall at a five percent false positive rate sets the threshold on the same months it reports. A bank has to fix its threshold before the applications arrive, and the false positive rate it then gets is whatever the new months give it. | The report and the page publish both: the suite's measure, and the recall, false positive rate and predictive equality on months 6 and 7 at the threshold fixed on month 5. The recall gate uses the fixed threshold, and a gate watches how far the realised false positive rate drifts from its target. |
| 2026-09-27 | attrition | The first split trained on every month end up to December 2024 and tested on the summer of 2025. A December 2024 label says whether the customer left by December 2025, which the bank could not have known when it scored the summer of 2025. | Training uses month ends September 2023 to February 2024, validation March to May 2024, and test June to August 2025, so every training and validation outcome was known before the first test month end. The gap is stated in the report beside the split. |
| 2026-09-27 | attrition | Sorted by probability alone, the top of the retention list fills with new customers holding a few hundred dollars. A retention call costs the same whoever answers. | The list is ranked by expected balance at risk. The report publishes, for the top tenth of the test months, the share of closers and the share of their balances found by each ordering, so the trade between the two is visible. |
| 2026-09-27 | cure | A loan thirty days past due cures far more often than one ninety days past due. A model that only learned the bucket would beat the base rate comfortably and add nothing a collections supervisor does not already do by working the thirty day list first. | The bucket rule is a benchmark in the report, and a critical gate requires the model's test AUROC to beat the rule's by two hundredths. Cure rates by bucket are published beside the model's so the rule's strength is visible. |
| 2026-09-27 | cure | A loan that is past due in June and July 2025 appears in the last training month and the first validation month, with nearly the same features and the same outcome. Early stopping and the choice between the two candidates were being made on loans the fit had already seen. | One loan in five is held out of the fit by a hash of its identifier, and early stopping and the choice of champion use only held out loans in the validation months. The test months begin three months after validation ends, so every outcome used to fit or choose the model was known before the first test month end, and a check in the trainer refuses to run otherwise. |
| 2026-09-27 | relief | Complaints received in late 2024 ended in monetary relief about one time in seventy; by mid 2026 it was closer to one in fourteen. A model fitted on the early months will score the later ones too low, and a probability threshold chosen on 2025 would flag almost nothing in 2026. | The operating point flags the top twentieth of scores, set on the validation months immediately before the test months. Reliability on the test months is published with the drift stated beside it, the calibration gate is non critical, and the monitoring plan resets the operating point when the monthly base rate moves by half. |
| 2026-09-27 | relief | Consumers write about their age, their disability, their military service and their family. A text model can learn that an elderly widow is more likely to get a refund, which is a protected basis entering a routing decision by the back door. The Bureau's own tags mark older Americans and servicemembers. | `rules/proxies.yaml` now lists words a text model may not learn from, and the relief model drops them, and any phrase containing them, from its vocabulary. The Bureau's tags are never features; they are used only to audit flag rates by group, and a gate watches the ratio. |
Characteristic stability, scorecard
| Characteristic | Stability index, last six months against development |
|---|---|
| purpose | 0.003 |
| oldest_trade_months | 0.003 |
| revolving_utilization | 0.002 |
| bureau_score | 0.000 |
| dti | 0.000 |
| employment_months | 0.000 |
| loan_amount | 0.000 |
| inquiries_6m | 0.000 |
Source: generated, every installment application, approved or declined, seed 20260831, as of 2026-08-31.