Model risk · validation report · pd_challenger
Validation report: Probability of default challenger (monotone GBM)
Harborline Bank is fictional. This is a demonstration on generated data and public datasets; no real customer, account or transaction of any real institution appears here.
Nothing here is a credit decision, a fraud determination, a suspicious activity finding or investment advice.
| Model | pd_challenger |
| Purpose | Test whether a gradient boosted model beats the scorecard enough to justify its opacity |
| Tier | 1: Informs a decision a person makes, with a method whose behaviour is hard to inspect |
| Role | challenger |
| Status | validated |
| Owner | Head of analytics |
| Prediction time | The application date |
| Outcome | 90 days past due or charge off within 12 months of origination |
| Last validated | 2026-08-31 |
| Conclusion | not approved |
Purpose and scope
The challenger exists to answer one question for the model risk function: would a gradient boosted model decide installment loan applications better enough than the scorecard to justify giving up the scorecard's transparency? It is shadow scored on every application, including declines, and decides nothing.
Conceptual soundness
LightGBM on the same raw characteristics the scorecard considered, with monotone constraints on every numeric characteristic in the direction the scorecard's weight of evidence runs, so a higher bureau score can never raise predicted risk and a higher utilisation can never lower it. Categorical characteristics are unconstrained. Early stopped on the validation months, 132 trees. Without the constraints a tree model can learn non monotone effects that a lender could not defend to an examiner or explain to an applicant.
Data and assumptions
Identical to the champion's: the same booked loans, the same censoring, the same split by origination month, the same leakage test and the same exclusion of every protected attribute and proxy.
Development evidence
Development AUROC 0.776, validation 0.670. Its rank order agrees with the champion's on every application at a Spearman correlation of 0.962, so the two models mostly disagree at the margin rather than about who is risky.
Outcomes analysis on held out data
On the test sample the challenger's AUROC is 0.667, an improvement of 0.005 over the champion. Stability from development to test is 0.0035. By age band, at the threshold that approves the champion's share of applicants:
| Group | N | Event rate | Mean score | Calibration gap | Auroc | Approval rate | Adverse impact ratio |
|---|---|---|---|---|---|---|---|
| 18 to 24 | 492 | 6.71% | 3.90% | -2.8% | 0.625 | 18.3% | 0.91 |
| 25 to 34 | 611 | 7.04% | 4.33% | -2.7% | 0.663 | 19.8% | 0.99 |
| 35 to 49 | 1,238 | 6.87% | 4.08% | -2.8% | 0.708 | 18.1% | 0.90 |
| 50 to 64 | 988 | 6.98% | 3.81% | -3.2% | 0.650 | 20.0% | 1.00 |
| 65 and over | 413 | 5.81% | 3.83% | -2.0% | 0.651 | 19.9% | 0.99 |
Calibration on the test sample, ten equal count bins:
| Bin | Accounts | Predicted | Observed |
|---|---|---|---|
| 1 | 375 | 0.85% | 2.93% |
| 2 | 375 | 1.15% | 3.20% |
| 3 | 374 | 1.51% | 3.74% |
| 4 | 374 | 1.99% | 5.35% |
| 5 | 374 | 2.59% | 4.81% |
| 6 | 374 | 3.38% | 6.15% |
| 7 | 374 | 4.43% | 6.42% |
| 8 | 374 | 5.72% | 7.49% |
| 9 | 374 | 7.47% | 7.75% |
| 10 | 374 | 10.90% | 20.05% |
Source: generated, booked installment loans with an observed outcome, seed 20260831, as of 2026-08-31.
Benchmarking
Against the champion first, then the logistic model and the base rate, on the same test sample.
| Model | AUROC | Gini | Brier | ECE |
|---|---|---|---|---|
| This model | 0.667 | 0.334 | 0.0624 | 0.0279 |
| Champion scorecard | 0.662 | 0.324 | 0.0626 | 0.0286 |
| Logistic on raw characteristics | 0.668 | 0.336 | 0.0624 | 0.0285 |
| Base rate | 0.500 | 0.000 | 0.0641 | 0.0307 |
Source: generated, booked installment loans with an observed outcome, seed 20260831, as of 2026-08-31.
Limitations
- Cannot produce Regulation B reasons on its own; reasons would have to come from a separate explanation method, and a bank would have to validate that method too.
- Shares every limitation of the champion's data.
Ongoing monitoring plan
While it stays in shadow:
| Measure | Trigger | Action |
|---|---|---|
| Rank agreement with the champion | below 0.7 | Investigate before any promotion case |
Effective challenge
From DECISIONS.md:
| Date | Challenge | Response | What changed |
|---|---|---|---|
| 2026-09-26 | The first fit's development AUROC was far above its validation AUROC. | With fifteen leaves and a small minimum leaf, the trees were learning the development months' noise; early stopping alone did not close the gap. | Seven leaves and a minimum of 150 loans per leaf. The development to validation gap narrowed and the validation and test AUROC held. |
| 2026-09-26 | The challenger's test AUROC is above the champion's. | The gain has to pay for what the bank gives up. A gradient boosted model cannot state Regulation B reasons directly; declined applicants would still get reasons computed from the scorecard, which means explaining one model's decision with another model. A small gain does not buy that. | Promotion requires a test AUROC at least 0.01 above the champion's, stated as a gate in the validation report, together with the stability and age band gates. |
Promotion gates
| Gate | Threshold | Measured | Result |
|---|---|---|---|
| Test AUROC at least 0.01 above the champion | at least 0.01 | 0.0054 | FAIL |
| Score stability, development to test (PSI) | at most 0.25 | 0.0035 | pass |
| Lowest adverse impact ratio across age bands at the champion's approval rate | at least 0.8 | 0.9029 | pass |
| Expected calibration error on test | at most 0.02 | 0.0279 | FAIL |
| Rank agreement with the champion on every application (Spearman) | at least 0.7 | 0.9617 | pass |
Validation conclusion
Not approved for promotion. The challenger is at least as stable and as fair as the champion and ranks marginally better, but the improvement is under the materiality margin set before evaluation, and it would cost the bank directly stated adverse action reasons. It stays in shadow and is re-evaluated at the next validation.
not approved. Conditions:
| Condition |
|---|
| Test AUROC at least 0.01 above the champion (measured 0.00535) |
The integration contract
| Line | What this model does |
|---|---|
| trained | Same development sample as the champion, early stopped on the validation months. |
| timed | Same application record as the champion, under the same leakage test. |
| calibrated | Reported on the out of time test sample beside the champion's calibration. |
| useful | Judged only against the champion, on the same test sample, by a stated materiality margin. |
| fair | The same age band audit, at the threshold that approves the champion's share of applicants. |
| gated | Promotion requires every gate in this report; it has not been promoted. |
| served | Shadow scored in batch on every application, including declines; not served for decisions. |
| integrated | Its scores sit beside the champion's in the applications table for the comparison and nothing else. |
| monitored | Rank agreement with the champion and its own stability, monthly, while it remains in shadow. |
| documented | This validation report and docs/models.md. |
| bounded | Not used for any decision while in shadow; cannot produce Regulation B reasons without a separate explanation method. |
| validated | This report, rendered from the manifest. |
| explainable | Monotone constraints keep each characteristic's effect in the scorecard's direction; feature contributions are published, but adverse action reasons still come from the champion. |