Model risk · validation report · cure

Validation report: Cure within ninety days for delinquent loans

Harborline Bank is fictional. This is a demonstration on generated data and public datasets; no real customer, account or transaction of any real institution appears here.

Nothing here is a credit decision, a fraud determination, a suspicious activity finding or investment advice.

Model cure
Purpose Order the collections queue by the balance likely to be lost without a call: balance times the probability the loan does not cure
Tier 2: Informs a decision a person makes; moderate materiality
Role sole
Status in use
Owner Head of analytics
Prediction time Each month end, for loans thirty or more days past due; features use that month and earlier
Outcome The loan is current, or paid off, at one of the next three month ends
Last validated 2026-08-31
Conclusion approved

Purpose and scope

Order the collections queue by the balance the bank is likely to lose without a call: each past due loan's balance times the probability it does not return to current within ninety days. The model informs whom a collector calls first; it never changes terms or charges anything off, so it is tier two.

Out of scope: hardship programmes, settlement offers, and the decision to charge off.

Conceptual soundness

Whether a delinquent loan cures depends on how late it is, how long it has been late, whether the borrower paid anything last month, and the loan itself. Two candidates were fitted on the same features, a gradient boosted classifier and a logistic regression; the champion was chosen on held out loans in the validation months, and the boosted model had to beat the logistic by 0.005 of AUROC to earn its complexity. The model is the logistic regression (0.760 for LightGBM against 0.762 for the logistic regression on held out loans).

The benchmark that matters is the rule collectors use today, the bucket: each bucket's training cure rate used as a score.

Data and assumptions

Every loan thirty or more days past due at a month end. The outcome is whether it is current or paid off at one of the next three month ends, so it is known three months later, and the split respects that: training on month ends September 2023 to June 2025, validation on month ends July to September 2025, test on month ends December 2025 to May 2026 (8,104 delinquent loan months, 2,663 of them cured).

A leakage test recomputed the run of late months, the earlier delinquency episodes and last month's payment for 200 sampled loan months from the loan's own history and matched every one. No demographic attribute is a feature.

Development evidence

Development AUROC 0.784, validation 0.777. What moves the score most (standardised coefficients for a logistic regression, share of gain for a boosted model):

feature weight
months in a row past due -0.98
interest rate -0.61
a card account 0.29
ninety or more days past due -0.24
a mortgage -0.13
sixty days past due -0.09
earlier delinquencies -0.08
card utilisation 0.06
share of the payment made last month 0.06
made a payment last month 0.06
balance -0.05
age of the loan 0.04
unemployment 0.01

Outcomes analysis on held out data

On the test months the model reaches an AUROC of 0.787, against 0.761 for the bucket rule and 0.785 for the LightGBM.

By bucket, on the test months:

Group N Event rate Mean score Calibration gap Auroc
30 3,069 58.52% 57.40% -1.1% 0.598
60 1,845 30.08% 32.38% +2.3% 0.545
90+ 3,190 9.78% 10.40% +0.6% 0.638

Source: generated, Harborline installment, card and mortgage loans past due at a month end, September 2023 to August 2026, seed 20260831, as of 2026-08-31.

Calibration on the test sample, ten equal count bins:

Bin Accounts Predicted Observed
1 811 3.42% 4.07%
2 811 6.99% 8.63%
3 811 12.22% 11.84%
4 811 18.36% 16.28%
5 810 27.17% 26.79%
6 810 35.84% 32.35%
7 810 45.46% 45.31%
8 810 54.33% 53.95%
9 810 60.44% 60.25%
10 810 67.90% 69.26%

Source: generated, Harborline installment, card and mortgage loans past due at a month end, September 2023 to August 2026, seed 20260831, as of 2026-08-31.

Benchmarking

Against the base rate, the bucket rule and the candidate that was not chosen, on the test months:

Model AUROC Gini Brier ECE
This model 0.787 0.573 0.1718 0.0107
Base rate 0.500 0.000 0.2209 0.0266
The bucket rule 0.761 0.523 0.1753 0.0217
LightGBM on the same features 0.785 0.569 0.1723 0.0137

Source: generated, Harborline installment, card and mortgage loans past due at a month end, September 2023 to August 2026, seed 20260831, as of 2026-08-31.

Limitations

  • The generator decides cures from a stated process with the loan's latent risk in it. Real cures depend on things no loan file holds: a new job, a family emergency, a collector's tone.
  • A call changes the outcome it predicts. Once the queue is worked in this order, the realised cure rates are no longer a clean measure of the model.
  • Paid off counts as cured. It is rare from delinquency here, and a real bank might treat a payoff through refinancing elsewhere differently.

Ongoing monitoring plan

Monthly:

Measure Trigger Action
Realised cure rate by queue decile, monthly the top decile cures more often than the bottom Refit
Score PSI, monthly against validation above 0.25 Revalidate
The bucket rule's AUROC beside the model's the margin falls below the gate for a quarter Revalidate

Effective challenge

From DECISIONS.md:

Date Challenge Response What changed
2026-09-27 A loan thirty days past due cures far more often than one ninety days past due. A model that only learned the bucket would beat the base rate comfortably and add nothing a collections supervisor does not already do by working the thirty day list first. The base rate is the wrong benchmark for this model. The benchmark is the rule in use: each bucket's cure rate from the training months, applied as a score. The bucket rule is a benchmark in the report, and a critical gate requires the model's test AUROC to beat the rule's by two hundredths. Cure rates by bucket are published beside the model's so the rule's strength is visible.
2026-09-27 A loan that is past due in June and July 2025 appears in the last training month and the first validation month, with nearly the same features and the same outcome. Early stopping and the choice between the two candidates were being made on loans the fit had already seen. The attrition model memorised customers for the same reason. Month ends are not independent draws when the unit is a loan that stays delinquent for months. One loan in five is held out of the fit by a hash of its identifier, and early stopping and the choice of champion use only held out loans in the validation months. The test months begin three months after validation ends, so every outcome used to fit or choose the model was known before the first test month end, and a check in the trainer refuses to run otherwise.

Promotion gates

Gate Threshold Measured Result
Test AUROC above the bucket rule the collections team uses today at least 0.02 0.0251 pass
Test AUROC against the other candidate's (the choice on validation must not cost more than a hundredth) at least -0.01 0.0019 pass
Test AUROC at or above the floor at least 0.7 0.7866 pass
Point in time recompute of sampled delinquent loan months at least 200 200.0000 pass
Expected calibration error on test at most 0.05 0.0107 pass

Validation conclusion

The model orders the queue the way its purpose asks and beats the rule in use on months it never saw. It is approved for ordering the collections queue.

approved. Conditions:

Condition
None

The integration contract

Line What this model does
trained On month ends September 2023 to June 2025, whose three month outcomes had all closed before the test months began.
timed Every feature comes from the month end or earlier; a leakage test recomputes the run length, earlier episodes and last payment for 200 sampled loan months.
calibrated Reliability on the test months is published; the queue multiplies the balance by one minus the probability, so calibration matters.
useful Beats the base rate, the bucket rule collectors use today, and the other candidate on months it never saw.
fair No demographic enters; the features describe the loan and its payments. Cure rates by bucket are published beside the model's.
gated Every gate in this report, including the margin over the bucket rule.
served Scored monthly in the pipeline; the queue is published with the collections page.
integrated Orders the collections queue by balance at risk times the probability of not curing.
monitored Monthly: realised cure rate by queue decile, score PSI, and the rule's AUROC beside the model's.
documented This validation report and the collections page.
bounded Orders calls. Never changes terms, never charges off a loan, and never goes to a borrower as a score.
validated This report, rendered from the manifest before the queue uses the order.
explainable Each queued loan shows the two features that lowered its chance of curing most, in plain words.