Your credit model has a collections dependency
A credit model learns the performance borrowers produced under the lender's own servicing and collections policy, not a policy-free truth about the borrower. That is not automatically an error, but it can create model risk when the policy changes or varies opaquely across segments. A mechanism argument with a test, not a measured effect.
Two borrowers look the same at origination: same score, same product, same limit. Both miss a payment in month four. One is reached within days and offered a plan that fits their month, and cures. The other is contacted late, worked with a generic sequence, and charges off. A year later your data records one as good and one as bad, and that pair goes on to teach the next scorecard. The question worth asking is what, exactly, it taught.
Before the argument, its limit, because it belongs up front. This mechanism matters only if three conditions hold together.
- Post-origination servicing can materially affect the label before it is fixed.
- Treatment assignment or its effectiveness varies with information visible at origination.
- The model is later used under a different policy, or judged against a common reference policy.
We do not establish the prevalence or magnitude of these conditions here. If any one fails, the concern is small. The rest of this piece explains why they matter and gives a test for whether they hold on your book.
What the label actually measures
An application model predicts a label, some delinquency, default or loss event in a performance window, from features known at origination. That label is not a fact of nature about the borrower. A loan does not resolve on its own; the recorded result is co-produced by your servicing. It is the outcome realised under the policy the account actually experienced, call it Y(π). You observe Y(π) for the policy that account got, never a policy-free outcome. A model fit on that data predicts performance inside the system you ran, which is legitimate, and narrower than most teams assume.
Policy dependence, not corruption
It is tempting to say "so your labels are corrupted." Wrong word. If a channel, region or customer type performs worse under the collections policy you actually run, that is a correct prediction about your system as it operates. The model is validated under that historical policy; validity under another policy is a separate claim that must be demonstrated. This sits beside the selective-labels problem, not on top of it: reject inference concerns outcomes you never saw because you declined the applicant; this concerns observed outcomes conditional on the treatment each account received. A causal question, not a data-cleaning one.1
Three problems, kept separate
Policy dependence. The model learns realised outcomes under the historical regime. A property of the target, not a defect.
Policy transport. Change vendors, hardship rules, contact strategy or escalation thresholds, or the pricing and limits upstream, and a model trained under the old policy may no longer rank or calibrate correctly under the new one. This is the part with teeth.
Segment-correlated treatment. If servicing varies systematically with origination-visible features, the model can encode those treatment patterns. That may improve prediction under the same historical policy, reduce it under a changed policy, and raise a conduct question either way. A coefficient is not wrong merely because the feature acts through the lender's behaviour rather than the borrower's.
When it becomes model risk, and how it bites
Start with the cheap case. Where a lender ranks applicants and sets a cutoff, a servicing change that moved every outcome by a roughly constant amount would shift the level but preserve the ranking, and re-thresholding would approve almost the same people.
A level shift is cheap for accept-or-decline. It still matters directly for pricing, expected loss, limits and provisioning, which use the probability itself.
The expensive case is a change that alters who ranks above whom, which happens when the servicing effect is correlated with features the underwriter can see. The old model's ordering of those segments is fitted to the old policy, and re-thresholding alone cannot repair a material ranking change. That is the calibration-versus-discrimination distinction: miscalibration that preserves order is cheap for a cutoff and still costs you on anything priced off the probability; loss of ranking is what changes who you approve.
Most: the target includes post-delinquency outcomes (cure, charge-off, recovery, realised loss), treatment happens before the label is fixed, its assignment or effectiveness varies with origination-visible features, and the future regime differs from the historical one.
Little: the label is fixed before any material servicing can affect it, servicing is uniform and stable, and the model runs under the same policy that generated its labels. Establish that timing product by product, not from the label's name.
An illustration, and what it is not
To make the mechanism visible we built a simulation, and its limits come first: it is a synthetic existence proof, not a measurement, and no number in it estimates prevalence or magnitude on any real book. The design is the honest part. An explicit grid of training policy against evaluation policy, over three servicing policies: a reference policy with no shortfall, a uniform bounded shortfall, and a segment-correlated one mean-matched to the uniform in probability space, so the average outcome hit is equal and only its shape differs. We train the same model under each and compare models within a single evaluation policy: an in-policy model against a transported one, scored on the same outcomes.
Two things follow. A model trained under the segment-correlated policy is excellent under that same policy, but on the reference policy its ranking trails a model trained there, and a flexible model transports worse than a logistic one because it fits the treatment pattern harder; an oracle that knows each policy's probabilities sets the ceiling. Holding the other components fixed, the extra gap is attributable to the imposed treatment heterogeneity: remove it and the gap returns to zero within noise, which a test in the code checks explicitly.
The catch is what the toy grants itself. Its segment feature is irrelevant to repayment by construction, so the treatment artefact is cleanly separable from real risk. On a real book the same attributes carry genuine credit signal, and separating an operational effect from a real risk correlation in the same feature is the identification problem this cannot solve. It gestures at a mechanism; it does not measure your book.
If it is real, it compounds, and governance can miss it
Applicants a policy-shifted model would decline never become customers, never generate an outcome, and are absent from the next redevelopment, so a decision system can curate its own training data, a structure familiar from feedback loops elsewhere.2 The analogy is structural, not evidence of magnitude in lending. Routine governance can be blind to it: validation against outcomes from the same stable policy will not reveal that a model fails under a different reference policy, because the development and out-of-time samples were both produced under that policy and agree with each other, and that agreement is not evidence the model transports. Only a deliberately identified contrast between policies shows it. The timing can also be procyclical: a stressed vintage is serviced under capacity-constrained rules, those outcomes mature into a redevelopment that absorbs the stressed policy, and when the policy later normalises a transport mismatch is left behind, remembering that a low default rate alone cannot distinguish better selection from a more restrictive policy.
The regulatory picture, and it has just moved
The ground shifted in 2026, in ways that cut both directions. On model risk, SR 26-2 superseded SR 11-7 and SR 21-8 in April 2026; it is risk-based and non-prescriptive, emphasises intended use, data relevance and monitoring as conditions change, and does not mention collections, so treating a material servicing change as a monitoring trigger is our inference, not a stated requirement.3 On fair lending, the CFPB finalised a Regulation B rule removing the effects-test language and stating that ECOA does not authorise disparate-impact liability, effective 21 July 2026;4 the operative rule and the pending challenge in National Fair Housing Alliance v. CFPB are distinct, and this addresses only Regulation B and ECOA, not other regimes, so it draws no general "no exposure" conclusion.5 Adverse-action rules still require specific principal reasons that accurately describe the factors actually scored, and a coefficient partly shaped by historical servicing does not by itself make a stated reason untrue.6 In the UK, CONC 7 and the Consumer Duty require forbearance and due consideration for customers in arrears; the data-quality reading here is ours.7
What to do about it
The useful version is a diagnostic. Three questions locate the exposure: what policy generated this label; is treatment assignment or effectiveness predictable from origination-visible features; will the model run under the same policy that generated its training data? If the answers are "a stable one," "no," and "yes," deprioritise the mechanism, document it, and monitor for material policy or population change, remembering that a negative finding can also come from weak treatment tracking or low statistical power.
- Inventory the post-origination interventions that can affect each model's target.
- Record treatment, eligibility, timing, vendor, agent, offer and policy version alongside the outcome.
- Separate in-policy validation from validation under a future or reference policy.
- Revalidate material underwriting models after a major servicing change.
- Test for treatment-effect heterogeneity across origination-visible segments.
- Put servicing-policy lineage in the model documentation, with named owners across credit, servicing, data science and model risk.
And a test, because a claim like this should be falsifiable. It is staged. First, check whether servicing varies with origination-visible features after controlling for delinquency state and eligibility. Second, using randomisation where you can, estimate whether the effect of treatment is heterogeneous across those features, the real question. Third, define a specific reference policy and estimate outcomes under it. Fourth, compare models against that common policy on ranking, calibration, expected loss and decision regret. If they agree, this argument is overblown for your book; if they diverge, you have a transport exposure no feature audit would have surfaced. Be precise about observability: outcome impact for applicants who would be newly declined needs randomised marginal approvals, a valid holdout or a defensible quasi-experiment, and champion/challenger testing contributes to this rather than proving it.8
Disclosure: Audun builds servicing and collections systems, so we benefit if lenders treat servicing quality and data lineage as strategic. That is the honest conflict. The mechanism here is testable, its prevalence and magnitude must be established on each lender's own portfolio, and nothing above depends on our product being the answer.
Notes
- On selection bias, a neighbouring but distinct problem: Lakkaraju, Kleinberg, Leskovec, Ludwig & Mullainathan, "The Selective Labels Problem," KDD 2017. Reject inference concerns outcomes never observed because credit was declined; the dependence here concerns observed, treatment-conditional outcomes, a potential-outcomes problem, not the same identification problem.
- Cited only as a structural analogy, not as evidence of magnitude in lending: Ensign et al., "Runaway Feedback Loops in Predictive Policing," PMLR 2018; Perdomo et al., "Performative Prediction," ICML 2020.
- Federal Reserve, SR 26-2, "Revised Guidance on Model Risk Management" (17 April 2026), which supersedes SR 11-7 and SR 21-8. It is risk-based and non-prescriptive and does not address collections; applying it to a servicing change is our inference.
- CFPB, final rule amending Regulation B, "Equal Credit Opportunity Act (Regulation B)," 91 Fed. Reg. (Apr. 22, 2026), effective 21 July 2026, removing the effects-test language and stating the Bureau's position that ECOA does not authorise disparate-impact liability.
- Challenge to the rule: National Fair Housing Alliance v. Consumer Financial Protection Bureau (filed 2026). Whether ECOA itself authorises disparate-impact liability is disputed; the operative rule and the litigation are distinct.
- Regulation B, 12 CFR 1002.9 and its Official Interpretations: adverse-action reasons must be specific and must relate to and accurately describe the factors actually considered or scored. CFPB Circulars 2022-03 and 2023-03 were withdrawn in 2025; the regulation and commentary remain. Not legal advice.
- UK: FCA Handbook, CONC 7 (arrears, default and recovery, including forbearance and due consideration) and the Consumer Duty. Contextual summary, not legal advice.
- Champion/challenger treatment testing on the delinquent book is standard collections practice (documented by the major bureaux and vendors, e.g. Experian, FICO). It is an instrument that can contribute to the test above, not sufficient on its own.
Reproduce it: the simulation code, seeds and environment that generate the figure are public at github.com/audun-collection/audun-website-01/news/sim. Every value in the figure caption traces to that code.