Zum Inhalt springen
← All coverage

Coverage note · Credit Risk · Credit Risk · ML

CreditForge — Bank-Grade Credit Risk Platform

The 80% that makes a PD model bank-credible — leakage-safe targets, out-of-time validation, calibration, reason codes, fairness, and drift monitoring.

Personal projectHugging Face SpacesDemo live · live from Hugging Face

Key metrics

Risk chainPD, LGD, EAD to EL · Basel / IRB methodology
Validationout-of-time vintage split · Freddie Mac Single-Family sample
PD qualityGini/KS, isotonic calibration, PSI · CI-gated against thresholds

1. Out-of-time vintage split on the Freddie Mac Single-Family sample; PD validated by Gini/KS, isotonic calibration, and PSI stability — CI-gated against thresholds.

The problem

Expected Loss is one line — EL = PD × LGD × EAD. You can fit a PD model on a Kaggle dataset in an afternoon and feel like you have done credit risk. You have not. The model is roughly 20% of the work.

The other 80% is the scaffolding that makes a model a validation committee would actually sign off on. The silent killer is leakage: a random train/test split lets the model peek at the future and reports a discrimination score that evaporates the moment it sees next quarter's originations.

Architecture

  1. 01

    Bronze

    Raw Freddie Mac Single-Family loan + monthly-performance panels, vintage-partitioned.

  2. 02

    Silver

    Cleaned, point-in-time 12-month default flag, performance joined.

  3. 03

    Gold

    Leakage-safe feature matrix + forward target, vintage-tagged.

  4. 04

    Out-of-time split

    Train on old vintages, test on newer ones — never a random shuffle.

  5. 05

    PD models

    WoE scorecard (logistic, interpretable) + LightGBM challenger.

  6. 06

    Calibration

    Isotonic — turns a good ranking into a calibrated probability.

  7. 07

    LGD + EAD → EL

    Two-stage LGD and EAD/CCF combined into Expected Loss.

  8. 08

    Validation + governance

    Gini/KS/PSI, SHAP reason codes, fairness tests, model card; FastAPI serving + Next.js Risk Cockpit + drift monitoring.

Key tradeoffs

Out-of-time vintage split instead of a random split.

WhyThe only honest test of "will this work on loans I have not seen yet." A random split inflates a Gini that does not survive deployment.

Train a WoE scorecard and a LightGBM challenger side by side.

WhyThe scorecard is the interpretable model you defend in the room; the challenger measures how much signal the simple model leaves on the table.

Isotonic calibration as a first-class step.

WhyA credit decision needs a calibrated probability, not just a rank. A well-ordered but miscalibrated model prices every loan wrong.

Risk Copilot agents orchestrate; the classical ML core computes every number.

WhySHAP explains the drivers and the LLM narrates them — it never invents a Gini.

Eval results

Gini / KS
Discrimination

Out-of-time test on Freddie Mac vintages; gains/lift curves.

Isotonic + HL
Calibration

Reliability curve and Hosmer-Lemeshow — predicted vs realized default rate.

PSI / CSI
Stability

Population and characteristic drift vs the training distribution.

SHAP · reason codes · fairness
Governance

Per-applicant adverse-action reason codes, protected-group fairness testing, and a documented model card.

Production proof

The artifact that keeps the numbers honest: the eval harness and monitoring gates that run in CI, not a one-off notebook result.

CI performance gates

CI · PASSING
Discrimination (Gini)gated
Calibration errorgated
PSI stabilitygated + monitored

creditforge/eval gates fail the build on regression, and PSI drift is monitored on a schedule in production — the model cannot silently degrade.

A credit-risk stack that answers the only question that matters inside a bank — "would your model-validation committee sign this off?" — end to end, on real GSE mortgage data.

Request coverage

I am focused on finance AI: credit risk, RegTech, AML, and agentic investment research. Open to roles, mentorship, and collaborators in fintech, quant, and bank AI.