Model risk · validation report · pd_scorecard

Validation report: Probability of default scorecard (champion)

Harborline Bank is fictional. This is a demonstration on generated data and public datasets; no real customer, account or transaction of any real institution appears here.

Nothing here is a credit decision, a fraud determination, a suspicious activity finding or investment advice.

Model pd_scorecard
Purpose Rank and price installment loan applications by twelve month default risk and state the reasons for a decline
Tier 1: Material: informs a credit, fraud or regulatory decision directly
Role champion
Status in use
Owner Head of analytics
Prediction time The application date, before any loan exists
Outcome 90 days past due or charge off within 12 months of origination
Last validated 2026-08-31
Conclusion approved with conditions

Purpose and scope

The scorecard decides installment loan applications at Harborline: it ranks each application by the probability of default within 12 months, approves at or above a score of 580, and states the reasons for every decline. It also supplies the probability of default used in expected loss. It is the champion; pd_challenger is shadow scored beside it.

In scope: consumer installment loan applications through every channel. Out of scope: cards and mortgages, which it has never seen, and any use as a pricing model without a separate calibration.

Conceptual soundness

A weight of evidence scorecard is the method most consumer lenders still use for application decisions, for reasons that hold here: each characteristic's effect is monotone and visible in a table, the model is a logistic regression a reviewer can reproduce by hand, and the points give Regulation B reasons directly, as the characteristics where an applicant earned the fewest points against the best attainable. The alternatives considered were a logistic regression on raw characteristics, which ranks about as well but gives no natural reason codes, and a gradient boosted model, which is the challenger.

Each characteristic is fine classed into up to 20 quantile bins, then coarse classed until every bin holds at least 5% of applicants and weight of evidence is monotone. Characteristics enter only above an information value of 0.02 and below a weight of evidence correlation of 0.7 with any stronger one; a characteristic whose sign reverses in the joint fit is removed. Points are scaled to 600 at odds of 30 to one, with 20 points to double the odds. Of 13 candidates, 8 were kept:

characteristic iv kept
revolving_utilization 0.333 kept
bureau_score 0.310 kept
dti 0.233 kept
employment_months 0.074 kept
oldest_trade_months 0.069 kept
purpose 0.045 kept
inquiries_6m 0.030 kept
loan_amount 0.024 kept
annual_income 0.011 information value 0.0113 below the floor
housing_status 0.009 information value 0.0087 below the floor
term_months 0.008 information value 0.0084 below the floor
delinquencies_24m 0.004 information value 0.0036 below the floor
channel 0.002 information value 0.0021 below the floor

Data and assumptions

Generated data: installment loan applications to the fictional bank from its core banking generator, with the application record as it stood on the application date. The outcome is 90 days past due or charge off within 12 months of origination. Under the censoring calendar only loans originated at least 12 months before the as of date have an outcome; later ones are excluded, never counted as good. The split is by origination month: development 5,721 loans, validation 3,725, test 3,742. The validation months are the two vintages written under a loosened cutoff; the test months include the recession quarter.

Declined applicants have no outcome, so the scorecard is fitted on booked loans only. No reject inference was applied; the effect is that the scorecard's view of the lowest score bands rests on the applicants a legacy policy chose to book.

Age, ZIP code and county are held on the application for the fairness audit and are on the committed proxy list; none is a candidate characteristic.

Before any figure here was trusted, the pipeline recovered the generator's stated parameters:

parameter stated estimate interval_low interval_high tolerance within
Latent risk rank order (Spearman; perfect ordering is 1) 1.000 0.775 0.775 0.775 0.600 true
Latent risk coefficient 0.950 0.940 0.890 0.991 0.150 true
Macro sensitivity per point of unemployment 0.300 0.305 0.254 0.356 0.100 true
Loose vintage effect 0.450 0.470 0.397 0.543 0.150 true
Deposit beta, checking 0.020 0.020 0.020 0.020 0.010 true
Deposit beta, savings 0.380 0.380 0.380 0.380 0.030 true
Deposit beta, time 0.720 0.720 0.720 0.720 0.030 true
Recovery share at 24 months, installment 0.220 0.204 0.204 0.204 0.050 true

Source: generated, booked installment loans with an observed outcome, seed 20260831, as of 2026-08-31.

Development evidence

The fitted scorecard:

characteristic bin share bad_rate woe points
revolving_utilization (-inf, 0.0697] 5.0% 1.05% 1.331 96.1
revolving_utilization (0.0697, 0.1322] 10.0% 1.40% 1.038 91.3
revolving_utilization (0.1322, 0.1832] 10.0% 1.92% 0.715 86.1
revolving_utilization (0.1832, 0.3486] 30.0% 2.86% 0.307 79.4
revolving_utilization (0.3486, 0.48] 20.0% 4.29% -0.113 72.5
revolving_utilization (0.48, 0.5197] 5.0% 4.88% -0.249 70.3
revolving_utilization (0.5197, 0.6275] 10.0% 6.65% -0.578 64.9
revolving_utilization (0.6275, +inf] 10.0% 8.39% -0.829 60.9
bureau_score (-inf, 672] 5.3% 8.55% -0.850 60.9
bureau_score (672, 694] 15.4% 6.48% -0.549 65.7
bureau_score (694, 716] 19.7% 4.53% -0.170 71.7
bureau_score (716, 747] 24.9% 3.73% 0.032 74.9
bureau_score (747, 763] 10.4% 2.01% 0.666 84.9
bureau_score (763, 795] 14.4% 1.71% 0.835 87.6
bureau_score (795, +inf] 10.0% 1.23% 1.170 92.9
dti (-inf, 0.1193] 10.0% 1.22% 1.174 96.2
dti (0.1193, 0.1672] 10.0% 2.27% 0.544 84.5
dti (0.1672, 0.2117] 15.0% 2.56% 0.420 82.2
dti (0.2117, 0.2491] 15.0% 3.04% 0.244 78.9
dti (0.2491, 0.2713] 10.0% 3.32% 0.154 77.2
dti (0.2713, 0.2843] 5.0% 4.88% -0.249 69.8
dti (0.2843, +inf] 35.0% 5.95% -0.459 65.9
employment_months (-inf, 15] 5.9% 6.19% -0.502 62.7
employment_months (15, 25] 9.4% 5.21% -0.319 66.9
employment_months (25, 71] 40.1% 4.05% -0.055 73.1
employment_months (71, 206] 34.6% 3.39% 0.132 77.5
employment_months (206, +inf] 10.0% 1.92% 0.713 91.1
oldest_trade_months (-inf, 77] 10.3% 6.25% -0.511 69.3
oldest_trade_months (77, 95] 10.2% 4.78% -0.227 72.1
oldest_trade_months (95, 134] 24.9% 4.00% -0.042 74.0
oldest_trade_months (134, 150] 9.7% 3.79% 0.015 74.5
oldest_trade_months (150, 170] 10.3% 3.41% 0.126 75.6
oldest_trade_months (170, 196] 9.9% 3.00% 0.256 76.9
oldest_trade_months (196, +inf] 24.7% 2.83% 0.316 77.5
purpose home_improvement 15.8% 4.76% -0.223 67.7
purpose auto 25.7% 4.28% -0.111 71.1
purpose other 8.2% 3.82% 0.006 74.6
purpose debt_consolidation 37.5% 3.73% 0.031 75.3
purpose major_purchase 12.8% 2.19% 0.579 91.6
inquiries_6m (-inf, 1] 73.4% 3.45% 0.111 75.3
inquiries_6m (1, 2] 18.4% 4.76% -0.222 72.6
inquiries_6m (2, +inf] 8.3% 5.30% -0.335 71.7
loan_amount (-inf, 5,600] 10.1% 2.78% 0.335 83.2
loan_amount (5,600, 14,100] 55.1% 3.62% 0.063 76.1
loan_amount (14,100, 15,200] 5.1% 3.81% 0.011 74.7
loan_amount (15,200, +inf] 29.8% 4.63% -0.194 69.3

Source: generated, booked installment loans with an observed outcome, seed 20260831, as of 2026-08-31.

On the development sample the scorecard's AUROC is 0.722; on validation 0.666. Its scores move very little between development and test (population stability index 0.0032).

Outcomes analysis on held out data

On the out of time test sample (3,742 loans, 254 defaults) the scorecard reaches an AUROC of 0.662 (Gini 0.324, KS 0.238). At the cutoff it approves 80.3% of test applicants.

Calibration is where it falls short. The mean predicted default rate on the test months is 3.93% against an observed 6.79%, and on the validation months 4.44% against 9.61%. The ranking holds; the level does not, because the loosened vintages and the recession raised default for reasons no application characteristic carries. The known parameter recovery above attributes it: the vintage effect and the macro sensitivity are both present and both outside the scorecard.

By applicant age band on the test sample (age is never an input):

Group N Event rate Mean score Calibration gap Auroc Approval rate Adverse impact ratio
18 to 24 492 6.71% 3.83% -2.9% 0.590 82.3% 1.00
25 to 34 611 7.04% 4.20% -2.8% 0.658 78.4% 0.95
35 to 49 1,238 6.87% 4.00% -2.9% 0.708 78.8% 0.96
50 to 64 988 6.98% 3.76% -3.2% 0.652 81.7% 0.99
65 and over 413 5.81% 3.84% -2.0% 0.634 81.8% 0.99

The adverse impact ratio compares each band's approval rate with the most approved band.

Calibration on the test sample, ten equal count bins:

Bin Accounts Predicted Observed
1 375 0.54% 2.13%
2 375 1.03% 3.20%
3 374 1.52% 4.81%
4 374 2.09% 4.55%
5 374 2.76% 5.35%
6 374 3.53% 6.42%
7 374 4.42% 7.49%
8 374 5.56% 6.95%
9 374 7.11% 10.70%
10 374 10.75% 16.31%

Source: generated, booked installment loans with an observed outcome, seed 20260831, as of 2026-08-31.

Benchmarking

Against the base rate, a logistic regression on the same raw characteristics, and the gradient boosted challenger, all on the same test sample. The challenger's case for promotion is in its own report.

The same machinery on real data, 30,000 credit card accounts from the UCI Default of Credit Card Clients dataset (Yeh and Lien (2009), CC BY 4.0), split at random because the data has no dates:

measure generated uci
Loans or accounts with an outcome 13,188 30,000
Default rate 6.3% 22.1%
Characteristics kept 8 7
Scorecard AUROC, test 0.662 0.768
Scorecard Gini, test 0.324 0.536
Challenger AUROC, test 0.667 0.788
Logistic AUROC, test 0.668 0.750
Scorecard calibration error, test 0.0286 0.0125

The real accounts' scorecard:

characteristic bin share bad_rate woe points
pay_status_1 (-inf, -2] 9.3% 12.30% 0.706 91.9
pay_status_1 (-2, 0] 68.3% 14.12% 0.547 88.5
pay_status_1 (0, 1] 12.1% 34.71% -0.627 63.5
pay_status_1 (1, +inf] 10.4% 68.81% -2.050 33.1
pay_status_3 (-inf, -1] 33.2% 16.80% 0.341 81.0
pay_status_3 (-1, 0] 52.6% 17.27% 0.308 80.6
pay_status_3 (0, +inf] 14.2% 52.63% -1.364 60.3
pay_amt_1 (-inf, 0] 17.1% 36.32% -0.697 66.7
pay_amt_1 (0, 1,000] 8.6% 23.33% -0.069 75.8
pay_amt_1 (1,000, 2,000] 22.0% 22.63% -0.030 76.4
pay_amt_1 (2,000, 3,615] 17.2% 21.36% 0.045 77.5
pay_amt_1 (3,615, 4,370] 5.0% 20.00% 0.127 78.7
pay_amt_1 (4,370, 6,245] 10.0% 16.11% 0.391 82.5
pay_amt_1 (6,245, 8,015] 5.0% 14.78% 0.493 84.0
pay_amt_1 (8,015, 18,224] 10.0% 14.72% 0.498 84.1
pay_amt_1 (18,224, +inf] 5.0% 8.00% 1.184 94.1
limit_bal (-inf, 20,000] 8.3% 36.22% -0.693 68.1
limit_bal (20,000, 30,000] 5.4% 36.18% -0.691 68.1
limit_bal (30,000, 70,000] 17.3% 27.66% -0.298 73.1
limit_bal (70,000, 100,000] 10.6% 24.88% -0.154 74.9
limit_bal (100,000, 140,000] 9.2% 23.01% -0.051 76.2
limit_bal (140,000, 160,000] 5.9% 17.75% 0.275 80.3
limit_bal (160,000, 200,000] 11.2% 17.26% 0.309 80.8
limit_bal (200,000, 240,000] 8.8% 17.00% 0.327 81.0
limit_bal (240,000, 360,000] 15.2% 14.86% 0.487 83.0
limit_bal (360,000, +inf] 8.2% 11.08% 0.824 87.3
payment_ratio_avg (-inf, 0.03451] 9.7% 34.48% -0.617 71.8
payment_ratio_avg (0.03451, 0.07259] 33.8% 25.54% -0.189 75.3
payment_ratio_avg (0.07259, 0.1296] 9.6% 23.21% -0.063 76.3
payment_ratio_avg (0.1296, 0.3005] 9.6% 16.82% 0.340 79.7
payment_ratio_avg (0.3005, 0.6116] 9.7% 15.77% 0.416 80.3
payment_ratio_avg (0.6116, 1] 14.5% 14.93% 0.482 80.8
payment_ratio_avg (1, +inf] 9.6% 14.00% 0.556 81.4
payment_ratio_avg missing 3.5% 36.06% -0.686 71.2
bill_growth_6m (-inf, -0.0493] 20.0% 29.61% -0.393 72.9
bill_growth_6m (-0.0493, -0.02985] 5.0% 26.78% -0.253 74.3
bill_growth_6m (-0.02985, 0] 19.0% 23.52% -0.080 76.0
bill_growth_6m (0, 0.006685] 6.0% 19.11% 0.184 78.7
bill_growth_6m (0.006685, +inf] 50.0% 18.48% 0.225 79.1
utilization_1 (-inf, 0.2205] 45.0% 18.17% 0.246 77.0
utilization_1 (0.2205, 0.4258] 10.0% 20.40% 0.103 76.9
utilization_1 (0.4258, 0.526] 5.0% 23.22% -0.063 76.8
utilization_1 (0.526, 0.6302] 5.0% 25.44% -0.184 76.7
utilization_1 (0.6302, 0.9827] 25.0% 26.24% -0.226 76.7
utilization_1 (0.9827, 1.013] 5.0% 27.56% -0.292 76.6
utilization_1 (1.013, +inf] 5.0% 30.56% -0.438 76.5

Source: generated, booked installment loans with an observed outcome, seed 20260831, as of 2026-08-31.

Model AUROC Gini Brier ECE
This model 0.662 0.324 0.0626 0.0286
Base rate 0.500 0.000 0.0641 0.0307
Logistic on raw characteristics 0.668 0.336 0.0624 0.0285
Challenger (GBM) 0.667 0.334 0.0624 0.0279

Source: generated, booked installment loans with an observed outcome, seed 20260831, as of 2026-08-31.

Limitations

  • The data is generated. The scorecard proves the method and the controls, not a real portfolio's risk.
  • Fitted on booked loans only, with no reject inference.
  • Through the cycle by construction: it does not move with unemployment or with underwriting changes, which is why calibration fails on later vintages.
  • The UCI comparison uses a random split, so it cannot show drift, and the UCI data is a Taiwanese card portfolio two decades old, not a US installment book.
  • Simplified: no bureau attributes beyond the handful the generator produces.

Ongoing monitoring plan

Monthly, on every application and on each vintage as its outcomes mature:

Measure Trigger Action
Score PSI, monthly applications against development above 0.1 warns, above 0.25 breaches Investigate the shift; revalidate on breach
Characteristic stability index, each characteristic above 0.25 Review the characteristic's bins
Rolling test AUROC as outcomes mature falls 0.03 below the validation figure Redevelop
Observed over predicted default by vintage outside 0.8 to 1.25 for two vintages Recalibrate the intercept or apply an overlay

Effective challenge

Challenges raised during development and what each changed, from DECISIONS.md:

Date Challenge Response What changed
2026-09-26 On the validation months the scorecard's mean predicted default rate sat well under the observed rate, while its rank ordering held up. The validation months are the two vintages written under the loosened cutoff, and the test months carry the recession quarter. The hazard fit recovers a vintage effect close to the generator's stated one, and no application characteristic carries it, so a scorecard fitted on earlier vintages cannot see it. That is a calibration problem, not a ranking problem, and a bank would treat it the same way: keep the scorecard for decisions and adjust the probability used for loss. Calibration became a non critical gate, so the report concludes approved with a stated condition rather than approved, and expected loss in function 2 applies an observed to expected overlay by vintage instead of the raw scorecard probability.
2026-09-26 The fitted scorecard has no delinquency characteristic. Delinquencies in the last 24 months fell under the information value floor and were dropped. In this book the characteristic is nearly empty (most applicants have none) and carries little information once the bureau score is in the model, which already reflects payment history. Forcing it in would add a characteristic whose bins barely differ, and whose points would move with noise. The published information value table now lists every candidate with the reason it was kept or dropped, so a reviewer sees the decision instead of an absence, and the challenger, which sees every candidate, confirms the characteristic adds little.

Promotion gates

Gate Threshold Measured Result
Recovers the generator's latent risk ordering (Spearman) at least 0.4 0.7754 pass
Test AUROC at or above the floor at least 0.62 0.6618 pass
Test AUROC no more than 0.02 below the logistic benchmark at least -0.02 -0.0063 pass
Score stability, development to test (PSI) at most 0.25 0.0032 pass
Lowest adverse impact ratio across age bands at least 0.8 0.9524 pass
Point in time recompute of sampled rows at least 200 200.0000 pass
Expected calibration error on test at most 0.02 0.0286 FAIL

Validation conclusion

The scorecard ranks risk soundly out of time, recovers the generator's risk ordering, is stable, and passes the age band audit. It under predicts the level of default on the loosened and recession vintages. For decisions, which depend on rank, it is fit for use. For expected loss, its probability is used only with the vintage overlay in the lending function, and recalibration is the condition below.

The adverse action reasons it produced for declined test month applicants, by the first reason given:

reason declines share
Proportion of balances to credit limits on revolving accounts is too high 1,539 61.1%
Credit history reported by the credit bureau is insufficient or delinquent 680 27.0%
Income insufficient for amount of credit requested 284 11.3%
Length of employment 14 0.6%

approved with conditions. Conditions:

Condition
Remediate before the next validation: Expected calibration error on test (measured 0.02859)

The integration contract

Line What this model does
trained On loans originated in the first twelve months of the window, with outcomes observed under the censoring calendar.
timed Scores the application as submitted; a leakage test recomputes 200 sampled rows from the application record.
calibrated Points convert back to the logistic probability exactly; calibration on later vintages is reported, not assumed.
useful Benchmarked on the out of time test sample against the base rate, a logistic model on raw characteristics and the challenger.
fair Approval rates, AUROC and calibration by applicant age band; age is never an input.
gated Promoted only after every gate in its validation report passed or was accepted with a stated condition.
served POST /v1/scorecard/score in the API; in the browser from the published scorecard table on /studio.
integrated Sets the approve or decline decision at the stated cutoff and supplies expected loss in function 2.
monitored Score PSI and characteristic stability monthly; rolling AUROC and calibration as twelve month outcomes arrive.
documented This validation report, the scorecard table, and docs/models.md.
bounded Scores installment loan applications only; it has never seen a card or mortgage application.
validated This report, rendered from the manifest before the model is used anywhere.
explainable Every decline returns up to four adverse action reasons, the largest point shortfalls, in plain words.