Muskan Saha is a credit risk specialist — model developer, independent validator and quantitative analyst.
Muskan came to credit risk with a quant background, and she still reads a portfolio the way she reads an experiment: find the real distribution, question every assumption, and never trust a number that cannot be defended. That instinct is what model risk is actually about, and it shows in the work — she has developed the models behind credit decisions worth billions, from scorecards and PD term structures to loss forecasts and stress tests, and on separate engagements she has been the independent validator challenging them under IFRS 9, CECL, CCAR and IRB.
Her standard is models that are transparent, challenger-tested and built to survive the full credit cycle, not engineered to clear a review and fail when the macro turns — which is why she is as comfortable developing a model as she is defending it before a Model Risk committee.
Representative engagements, framed to respect client confidentiality — the approach and the standard, never the proprietary detail.
PPNR Forecasting · Tier-1 US Bank
Turned management judgment into a model the bank could defend.
She built a pre-provision net revenue forecast in Python, converting senior management's macro assumptions into quantifiable, auditable time-series drivers. She documented it end to end, stress-tested it against alternative scenarios, and presented the design, assumptions and results to the bank's Model Risk function.
PythonTime-seriesCCARModel governance
IRB Scorecard Validation · Retail Credit
Independent validation that underwriting could act on.
She validated application scorecards for credit cards and personal loans — discrimination (ROC-AUC, Gini), calibration and reject-inference bias — and translated the findings into guidance underwriting teams could act on, from cut-off calibration to risk segmentation.
IRBScorecardsSR 11-7Reject inference
IFRS 9 Term-Structure PDs · Deloitte
Lifetime credit losses, built to survive the macro cycle.
She derived 12-month and lifetime PDs from flow-rate analysis with forward-looking macro overlays, and benchmarked judgmental scorecards against gradient-boosted challengers, with SHAP-based explainability documented for independent review.
IFRS 9Flow-rateXGBoostExplainability
Model Risk Tools
Tools built outside client engagements — addressing gaps encountered in day-to-day model risk practice.
Open source · IFRS 9 · Disclosure benchmarking
IFRS 9 / ECL Peer Benchmarking
Turns ten banks' annual reports into a cited, analyst-ready IFRS 9 benchmarking set — without inventing a single figure.
Annual reports run to roughly 300 pages each, and the IFRS 9 disclosures an analyst actually needs are scattered through them. This tool ingests up to ten firms' PDFs and extracts five subsections per firm — core ECL and coverage, the staging table, impairment movement, model design (scenarios and weights, SICR, 30/90-DPD backstops, PMAs and overlays, climate, PD·LGD·EAD) and disclosure notes — then produces a browsable site, a consolidated PDF and a Power BI-ready master Excel.
The engineering is in the guardrails. Text is extracted layout-preserving, so multi-column tables don't collapse into ambiguous number streams. Retrieval narrows ~500 pages to the ~9 most relevant per section before anything reaches the model. A numeric value is kept only if its digits literally appear on a supplied page, and its citation is re-pointed to the page that actually contains it. Every figure carries a printed-page citation and a status flag — EXACT, PROXY, DERIVED or NOT DISCLOSED — and anything absent is listed rather than invented. Currency and units are detected per firm, and the approach was validated against five real bank annual reports. The dashboard below is shown with illustrative demo data for fictional peer banks — a live run extracts from real annual reports.
IFRS 9ECLCitation-grounded extractionPower BI
Peer comparisonFirm insightsStaging mix & coverageYear-on-year trendMovement & model designData quality & disclosure status
A working model-risk register that tiers every model, schedules its revalidation, and tracks findings through to closure.
Most model inventories are maintained in spreadsheets. This is a purpose-built alternative. Every model is scored on five factors — materiality, complexity, reliance, regulatory impact and uncertainty — into a weighted Tier 1/2/3 rating, and that tier drives the revalidation clock, flagging each model overdue, due soon or current. Findings are logged against models and tracked to resolution, with a dashboard over the top. It ships with a written methodology, CSV import/export for the Excel workflow teams actually use, and a Docker path for internal hosting — seeded with a fictional bank inventory for demonstration.
A validation framework for the models traditional model risk was never designed for — machine learning and LLMs, assessed side by side.
Most validation frameworks assume scorecards and regression. This one sets traditional ML and generative AI next to each other so the difference in how each must be validated becomes obvious, anchored to SR 11-7, the NIST AI Risk Management Framework and EU AI Act risk tiers. Models are rated on the dimensions that actually make AI risky — explainability, autonomy, output variability, drift, fairness, and for GenAI, hallucination and prompt-injection exposure — rolling transparently into a Tier 1/2/3 rating that sets revalidation cadence. Track A runs the traditional battery: discrimination (AUC, Gini, KS), calibration, PSI drift, feature importance and four-fifths fairness testing. Track B stress-tests a customer-facing banking chatbot through an editable system prompt and an impartial LLM-as-judge — weaken a guardrail and a test visibly fails. It runs fully simulated by default, or live against OpenAI, Azure OpenAI or any OpenAI-compatible internal gateway. Results are illustrative rather than production-calibrated, and the app says so.
SR 11-7NIST AI RMFEU AI ActLLM evaluation
Portfolio dashboardModel inventoryTrack A — traditional MLTrack B — generative / LLM
Model risk isn't a compliance tax — done right, it's how an institution earns the confidence to lend. Four positions that inform her practice:
IFRS 9 staging is more sensitive to model specification than to the macro path it tracks — the cliff-edge rarely falls where the threshold model places it.
A challenger model is not a benchmarking exercise. SR 11-7 is asking whether the production model can lose an argument.
SHAP built to satisfy a committee and SHAP that actually explains score variation are not the same deliverable — and the difference shows under stress.
GenAI earns its place in model documentation, anomaly flagging and challenger scaffolding; applied to model decisions themselves, the auditability problem remains unsolved.
Practitioners who see it differently are welcome to make the case.
Get in touch.
Muskan Saha is open to senior credit-risk and model-risk roles, welcomes conversations with practitioners working the same problems, and is happy to discuss any of the model risk tools above — including access or a walkthrough. Use the form below; only an email address is needed to reply.