COUNTERPARTY SCORECARD
Grade A–F methodology
The counterparty Grade A–F is a transparent, weighted composite of documented signals — never a black box. This page pins the exact weights, transforms, and letter bands to a semantic version, and backtests the HMDA-derived core across every historical activity-year transition so you can see whether the grade actually separated forward counterparty outcomes. The backtest numbers below are computed live from the panel, not hand-picked.
What changed in v1.1.0, and whether it helped
v1.1.0 makes one swap, on evidence. Pricing power — where a lender sat on HMDA rate spread — was a weighted component at 0.15 and is removed: the transform used the percentile as the sub-score, so the published grade rewarded a lender for charging a wider spread over APOR, and rate spread is the field regulators use to identify higher-priced lending. Growth volatility takes exactly that 0.15 weight, so the weight budget and every other component's relative influence are unchanged.
Discrimination is the probability that a randomly chosen counterparty-year that went on to distress scored BELOW one that did not; 0.5 is a coin flip. Removing pricing power and adopting growth volatility moved it +0.0646. The grade did not degrade. The v1.0.0 row is a fixed measurement taken on 2026-07-27 over the same panel; the v1.1.0 row is recomputed on every request.
Backtest — did the grade separate outcomes?
Each material counterparty (≥250 originations) is graded as of a year T using the exact production scorer on its two panel components — mix stability and credit posture — then we observe its outcome the next year (T+1). Window: 2019 → 2025, 13,583 graded counterparty-years.
Survival = still originating at T+1. Retention = median 1-year change in origination volume among survivors. Distress = exited the panel or lost more than half its volume. A monotonic drop in survival (and rise in distress) from A to F is the grade doing its job. Reported honestly: if a band is thin or the separation is modest, the numbers say so.
Letter-grade bands
Component weights & transforms
The composite is a weight-renormalized average over whichever components are AVAILABLE for a given counterparty — an unlinked seller is graded on its HMDA components alone, never penalized for missing Ginnie data. Components flagged in backtest are computable for the full historical panel; the rest enrich the live grade where their data links.
100 − 250 × total-variation drift of the conv/FHA/VA/USDA mix, first→as-of year (clamped 0–100)
Source: HMDA lender_annual_summary (product counts) · exists for 100.00% of graded counterparties · discrimination alone 0.5845
100 if denial ≤ market, else 100 − 400 × (denial − market) (clamped 0–100)
Source: HMDA denial rate vs market_national_year · exists for 100.00% of graded counterparties · discrimination alone 0.4655
100 − 2000 × 90+ delinquency rate on the linked serviced book (clamped 0–100)
Source: Ginnie Mae issuer disclosure (90+ DQ) · exists for 6.71% of graded counterparties
100 − 25 × max(0, cohort z-score of the conditional buyout rate) (clamped 0–100)
Source: Ginnie issuer speeds (cohort z of CBR) · exists for 4.07% of graded counterparties
100 − 100 × σ of year-over-year LOG volume growth across consecutive years (clamped 0–100); null under two growth observations
Source: HMDA lender_annual_summary (origination volume by year) · exists for 95.08% of graded counterparties · discrimination alone 0.6223
100 if the active-state count held or grew, else penalized by the share of states exited
Source: HMDA lender_state_year (active states) · exists for 100.00% of graded counterparties
Coverage figures were measured on 2026-07-27 over 13,583 graded counterparty-years / 3,173 companies. Credit posture is published at 0.4655 on purpose. Alone, it currently anti-predicts on this panel — the counterparties it marks down survived slightly better — and it is recorded here rather than quietly re-weighted, because a denial-rate weight is a fair-lending-adjacent decision that needs an owner, not a side effect of this release.
Components we evaluated and did NOT adopt
Six additions were proposed. One was adopted. The other five are published here with the numbers that rejected them, because a methodology that records only its acceptances invites the same five proposals every cycle. A candidate is weighted only if it exists for at least 4.07% of graded counterparties — the coverage of the weakest component already weighted — and discriminates at 0.55 or better on its own, and improves the composite on the subset it covers. Coverage alone is not a licence: a component that covers everything and predicts nothing only dilutes the grade.
Year-over-year origination volatility. Sharp growth is the classic precursor to credit and ops failure, and mix_stability measures product DRIFT, not SCALE change — a lender can hold a steady mix while tripling volume.
Coverage
95.08%
of graded entities
Samples
81.18%
of graded years
Dispersion
21.1
σ of sub-score
At floor
2.10%
pinned at 0
Alone
0.6223
discrimination
Incremental
+0.0566
0.5476 → 0.6042
The strongest single signal in the candidate set and stronger than either existing panel component; it holds in EVERY year transition it can be computed for (+0.034 to +0.065), including the low-distress years, so it is not a 2022 artefact. Adopted at 0.15 — the weight vacated by pricing_power.
Source: lender_annual_summary.origination_volume_dollars
A Herfindahl index over lender_state_year: a lender with 80% of its book in one state carries correlated house-price and regulatory exposure the current components do not measure.
Coverage
100.00%
of graded entities
Samples
99.99%
of graded years
Dispersion
32.4
σ of sub-score
At floor
15.10%
pinned at 0
Alone
0.4207
discrimination
Incremental
−0.0472
0.5517 → 0.5045
THE PROPOSED SIGN IS REFUTED. Concentration was PROTECTIVE across 2019→2024: penalising it costs 0.047 of discrimination, the largest degradation any candidate produced. Inverting the sign would gain 0.031, but that would publish 'we reward geographic concentration' on the strength of one rate cycle in which diversified national refi shops contracted hardest — the mechanism, not the coefficient, is what is unproven. Displayed, not weighted, in either direction.
Source: lender_state_year (origination counts by state)
Complaints normalised to origination dollars as a fifth counterparty-score component, decided on the backtest rather than on taste.
Coverage
13.80%
of graded entities
Samples
11.07%
of graded years
Dispersion
28.3
σ of sub-score
At floor
7.00%
pinned at 0
Alone
0.5271
discrimination
Incremental
+0.0021
0.5913 → 0.5934
Clears the coverage bar and both of the owner's stated hazards are DISPROVED — per-$B removes the size correlation (r = −0.025 with volume) and it does not double-count servicing quality (r = −0.023 with 90+ DQ) — but it carries almost no independent signal: 0.5271 alone and +0.002 incremental. Weighting it would add legal surface for no measured gain. It renders on the reputation tab instead, where the count is the product.
Source: cfpb_mortgage_complaints joined to the entity spine
A loan in forbearance is not counted as 90+ delinquent, so a servicer leaning on forbearance looks better on DQ than its book warrants — genuinely new information rather than a re-weighting.
Coverage
5.74%
of graded entities
Samples
5.18%
of graded years
Dispersion
34.3
σ of sub-score
At floor
14.60%
pinned at 0
Alone
0.4958
discrimination
Incremental
−0.0520
0.5926 → 0.5406
Clears the coverage bar and fails everything else: 0.4958 alone is a coin flip, and adding it COSTS 0.052 of discrimination on its own covered subset. The 'it is not a re-weighting' premise is also only half true — forbearance correlates with the 90+ DQ it is supposed to be independent of at r = 0.385.
Source: ginnie_issuer_forbearance over ginnie_issuer_monthly active loans
An issuer whose portfolio turns over rapidly is a different counterparty from one with a stable book at the same UPB and DQ.
Coverage
1.42%
of graded entities
Samples
0.39%
of graded years
Dispersion
41.9
σ of sub-score
At floor
28.30%
pinned at 0
Alone
0.6304
discrimination
Incremental
+0.0404
0.7422 → 0.7826
BLOCKED ON COVERAGE, not on merit: 45 of 3,173 graded entities (1.42%), below the 4.07% bar, on 53 samples containing 7 distressed years. The apparent +0.040 improvement rests on those 7 events and is not evidence. The double-count worry was checked and cleared (r = −0.043 with buyout intensity). Re-evaluate when the transfer file covers more than 2023-06 onward.
Source: ginnie_servicing_transfers UPB out / ginnie_issuer_monthly book UPB
HUD's own early-warning measure — the most externally credible component available, because it is not our arithmetic.
Coverage
0.66%
of graded entities
Samples
0.93%
of graded years
Dispersion
23.6
σ of sub-score
At floor
4.80%
pinned at 0
Alone
0.4477
discrimination
Incremental
−0.0522
0.7195 → 0.6673
The backlog sequenced this behind DD-042's all-lenders Early Warnings pull and that sequencing is confirmed: 21 of 3,173 graded entities (0.66%), every row at a single performance period, 47.6% of them pinned at the ceiling. It also anti-predicts at this sample size (0.4477) and costs 0.052 when added. Re-evaluate after DD-042 — the verdict here is about coverage, not about HUD.
Source: fha_de_performance.compare_ratio_total
Explicit non-goals
Two inputs that look obvious and are deliberately not components. Each is listed with the evidence in its favour first — a non-goal recorded only with its downside is a preference, not an argument.
What the evidence says for it: It was a weighted component at 0.15 through v1.0.0 and it does carry signal: 0.5812 alone over 82.68% of graded samples, worth +0.049 of composite discrimination. The reason it predicts is visible in the data — distress falls monotonically from 13.5% in the lowest spread decile to 6.9% in the highest, and mean book size falls with it. It selects for a business model (smaller, government/purchase-weighted books that had no refi volume to lose in 2022), not for counterparty quality.
Why it is a non-goal anyway: REMOVED in v1.1.0. The shipped transform used the percentile AS the sub-score, so a wider spread over APOR scored HIGHER — the published grade rewarded a lender for charging its borrowers more. Rate spread is the field regulators use to identify higher-priced lending, and a lender grade built on it sits directly on fair-lending exposure. Its 0.15 weight moved to growth volatility, which is measured on the same panel, covers 95.08% of entities instead of 80.84%, and scores higher on the same subset (0.6185 vs 0.6077). Rate spread remains on the scorecard as PEER CONTEXT and on /pricing-power as its own surface — the non-goal is using it as a GRADE COMPONENT.
Source: hmda_rate_spread via the pricing_power_lenders RPC
What the evidence says for it: It is the agency's own issuer scorecard, externally produced, and it spans more history than most of our Ginnie inputs.
Why it is a non-goal anyway: NOT ADOPTED. It is built on inputs that overlap ours and it already carries a copied z_cpr_3m, so folding it in makes our grade partly derivative of Ginnie's and double-counts servicing — which is weighted at 0.30 here. A grade that quietly re-uses another grade cannot be audited against its own components, which is the whole promise of this page.
Source: mortradar_issuer_iopp (201312–202605)
FREEThe grade is a computed, descriptive signal over public data — not investment advice, a credit rating, or a recommendation. Servicing, buyout and footprint components are not available for every lender in every past year, so they sit outside the panel backtest (documented, not fabricated). Source: HMDA public loan-application data (CFPB/FFIEC) + market_national_year. Back to counterparty scorecards.