Every Organization Needs a Risk Atlas.Most Don’t Have One for AI.
- Tejasvi A
- Aug 2
- 8 min read
What an AI risk function actually needs and what most teams are building instead.
I observed AI governance programs at financial institutions across India and rest of Asia over the past two years. In that time, I’ve asked the same question to almost every risk officer I’ve encountered: which of your deployed AI systems carries the highest residual risk right now and after your controls are applied? Not in general terms. Specifically, by system, by risk domain, with a number.
Most cannot answer it.
They can show me which models have been validated. They can produce a list of AI initiatives. They can show me a policy document and a framework crosswalk that took six months to build. What they cannot show me is a structured picture of what could go wrong across their AI portfolio, what stands between them and that outcome, and whether what stands there is actually working.
That gap is what a risk atlas closes. Not a model inventory. Not a validation report. Not a policy. A structured, living, measurable catalogue of failure modes — linked to controls, linked to owners, linked to evidence.
This is what one looks like, how to build it, and why the sequence most teams use is backwards.
01 — ANATOMY
What a Risk Atlas Actually Is
The term gets used loosely, so let me be specific about what I mean and what I don’t.
A risk atlas is not a model inventory. That tracks which models are deployed and their validation status. A risk atlas tracks what could go wrong with those models and what you’ve done about it. The difference sounds minor. In practice, a bank can have a complete model inventory and zero visibility into its AI risk exposure. I’ve seen this.
The organizing principle is risk domain, not product line. This matters more than it sounds. When you build a risk register per business division “the credit agent register,” “the fraud & AML detection register” you end up cataloguing hallucination risk in twelve separate places and missing it as a systemic issue. When you organize by domain (Model Accuracy & Reliability, Security & Robustness, Fairness & Bias, Governance & Accountability, Privacy & Data, and for agentic systems, Agent Autonomy, Multi-Agent Security, Tool & Affordance Risk), cross-cutting risks become visible. They can’t hide inside a product boundary.
Within each domain, every risk needs three things connected to it: the regulatory and industry frameworks it maps to (NIST AI RMF, OWASP LLM Top 10, MITRE ATLAS, ISO 42001, RBI, SEBI, DPDP); the deployment patterns where it’s relevant (RAG, autonomous agents, multi-agent, summarization, classification); and the controls that address it, with named owners and test evidence not policy references.
That last part is where most programs stop short. A control without a test result is a claim. In a governance context, claims are not evidence. I’ll come back to this.
And finally: residual scoring. After your controls are applied, what is the actual remaining exposure?
This is the number that goes to the board. An atlas without it is a compliance document. With it, it becomes something you can actually manage from.
The risk you have not named is a risk you cannot govern. This is not a philosophical point — it’s the most consistent finding in AI incident post-mortems. The failure mode existed in the literature. It was not in the catalogue.
02 — MATURITY
The Maturity Model — and Where Most Teams Actually Are
There’s something I want to flag before showing this model: most institutions I’ve worked with would assess themselves at Level 3. The evidence usually puts them at Level 2.
The difference is straightforward. Level 2 means you have a list. Level 3 means your list is connected to frameworks, to controls, to owners, to test evidence. The gap between those two things is enormous, and teams routinely underestimate it. They have a framework crosswalk that lived in a slide deck and was never operationalized. They have control assignments that haven’t been reviewed since the system was deployed. They call it Level 3 because they did the mapping work. They’re at Level 2 because nothing downstream of that mapping is functioning.
“Most institutions I’ve worked with would assess themselves at Level 3. The evidence usually puts them at Level 2.”
The five levels below describe a progression in what’s actually measurable and evidenceable, not just what’s been claimed.

I’ll be honest about Level 5: I’m not sure most institutions will reach it within a realistic planning horizon. The continuous monitoring requirement alone is a multi-year infrastructure investment. But that’s not the right frame for this model. The question isn’t whether you’ll reach Level 5 — it’s whether you know which level you’re at now, and what specifically would move you to the next one.
That question, applied honestly, is more useful than any target-state assessment I’ve seen.
03 — EVALUATION
How to Score Risk Severity Without Fooling Yourself
The mechanics here are standard: impact times likelihood, on a 4×4 matrix. What’s less standard is how teams misapply it.
The most common mistake I see: severity scores get used to prioritize compliance work rather than actual risk reduction. A risk scores Critical on the matrix but because it doesn’t appear in any framework the team has mapped, it gets deprioritized. A risk scores Medium but because it’s in NIST AI RMF, it gets treated as High. Neither of those is correct. The matrix drives prioritization. Frameworks provide the reporting vocabulary for findings. These are different jobs, and conflating them produces a risk program that is simultaneously over-resourced on low-impact compliance work and underinvested in genuine exposure.
Impact in a BFSI context has four dimensions: financial harm (direct loss, penalty, mis-selling liability), reputational harm, regulatory consequence, and operational disruption. Score each 1–4; the highest sets impact. Likelihood considers deployment frequency, capability maturity, and whether the system is customer-facing. Score 1–4 on the same scale.

One thing I’ve started doing in reviews: I ask teams to score five risks without referencing any framework just impact and likelihood on the matrix. Then I ask them to score the same five risks while looking at a framework crosswalk. The scores almost always shift upward in the second pass, even for risks where the framework mapping doesn’t tell you anything new about the actual harm. That shift is the compliance bias I’m describing. Worth running this exercise on your own team.
04 — CONTROLS
Measuring Control Strength and Why Most Scores Are Wrong
Here is the thing I want to say plainly, because I haven’t seen it written clearly anywhere else: most residual risk scores I’ve reviewed are fiction.
I have sat in board presentations where a risk heat map showed mostly yellows and greens comfortable numbers, comfortable conclusions. Those scores were built on control effectiveness assumptions assigned at design stage. Before a single test was run. The controls existed in architecture documents. They were in the GRC tool. They had owners. None of that means they worked.
Six months after one such presentation, an external audit found three of those controls non-functional in production. The board had been looking at a projection of intended behavior, not a measurement of actual behavior.
Control strength has four dimensions that need to be evaluated separately:
Coverage
Does the control apply to all of the risk surface, or only part of it? A hallucination guardrail that covers chat outputs but not API responses has partial coverage. This is the most commonly overestimated dimension. Controls get implemented for a specific deployment and never extended.
Key question: For what fraction of the risk surface is this control actually active?
Effectiveness
How much does the control actually reduce risk when it fires? A logging control is not the same as a blocking control. A classifier with 70% recall leaves 30% of the failure mode unaddressed. Effectiveness is an empirical question. You cannot answer it from a design document.
Key question: By what percentage does this control reduce the probability or impact of the risk?
Testability
Can you produce evidence? A policy document is a control. An automated test suite that runs daily against adversarial inputs is also a control. One produces evidence. One produces an argument. Your regulator will know the difference immediately.
Key question: What evidence exists that this control functioned as intended in the last 90 days?
Regulatory Alignment
Key question: Which specific clause or requirement does this control satisfy, and is that documented?

The fix I now recommend to every team I work with: score untested controls at zero. Not 40% for “we plan to implement this.” Not 20% “planning credit.” Zero, until there is test evidence. The discomfort of seeing a heat map full of reds is the correct response to not having tested your controls. It’s not a problem with the model it’s the model working. That discomfort should create urgency to test, not incentive to inflate the effectiveness assumption.
The residual score is what goes to the board. A domain heat map based on honest control effectiveness is the most actionable thing a mature risk function can produce. It tells leadership exactly where the next investment in controls will have the most impact.
05 — PRACTICE
Where to Start and Why the Usual Sequence Is Wrong
The standard advice is to start with a framework crosswalk. Map your risks to NIST AI RMF or OWASP LLM Top 10, get your alignment documented, and build from there.
I think this is the wrong first step. Frameworks are useful for regulatory reporting and audit alignment. They’re a subset of the actual risk landscape and in fast-moving areas like agentic AI, they lag deployment reality by 12–18 months. If you build your taxonomy by starting from NIST and filling it in, you’ll have good framework coverage and blind spots in the areas evolving fastest.
Start with the domain taxonomy. Then map to frameworks. The sequence matters.
→ Build the domain taxonomy first
Eight to twelve domains, organized by failure type — not by product, not by framework. This structure will outlast any individual risk entry or regulatory update. The taxonomy is the skeleton. Individual risks are content. If you start with content and add structure later, you’ll restructure it multiple times.
→ Tag by capability pattern from day one
Every risk entry needs tags for the deployment patterns where it’s relevant: RAG, autonomous agent, multi-agent, classification, chat. Retrofitting these tags to 120 entries is months of work you won’t want to do. An atlas that requires reading every entry to answer “what risks apply to my document extraction pipeline?” is an atlas no one uses.
→ Map to one framework before adding a second
Complete the NIST AI RMF crosswalk before starting OWASP, MITRE, or ISO 42001. Partial mappings across five frameworks are harder to maintain and less useful to an auditor than a complete, verified mapping to one. Add frameworks once the first is operational, not in parallel.
→ Score untested controls at zero
Already said this, but it bears repeating as an operational instruction. Design-stage controls get zero effectiveness credit. The discomfort of accurate numbers is a feature, not a problem. Accurate reds drive testing urgency. Inflated yellows drive false confidence.
→ Named individuals, not teams, on every risk
“The AI team” owns no risk. A named person, with a defined acceptable residual threshold and a review date, owns a risk. This is the single change that most reliably converts a risk catalogue from a compliance document into something people actually act on.
→ Review cadence should match deployment velocity
A bank deploying a new AI system every month needs monthly atlas reviews — or a deployment gate that asks: does this deployment introduce any risks not currently in the catalogue? An annual review cycle made sense when AI deployments were annual. It doesn’t now.

CLOSING
The Banks That Build This Now Will Have an Answer. The Others Will Have a List of Initiatives.
When a regulator walks in and asks “what is your AI risk exposure?” — a structured document with inherent scores, control test results, and residual ratings is a different kind of answer than a verbal summary of recent incidents and a slide about programs underway. Both take preparation. Only one of them is credible under scrutiny.
The gap between what most banks have and what they need here is not a large technical lift. It’s mostly a sequencing problem and a discipline problem. The sequencing: taxonomy before frameworks, frameworks before controls, controls before residual scoring. The discipline: not inflating effectiveness scores before you have evidence.
I’ve watched institutions build this right in four months. I’ve also watched institutions spend two years on a crosswalk that never got past the slide deck. The difference wasn’t resources. It was clarity about what the output was supposed to be and who was going to use it.
tejasviaddagada.com · AI Risk Governance · Banking & FSI
.png)


Comments