Stop Chasing Evidence to Start Informing Decision-Making

Walk into most GRC departments in large global companies and you'll find talented professionals spending their weeks on the same treadmill: updating a register, chasing a control owner for evidence, formatting a report nobody outside compliance will read. That work isn't worthless, audits need it and regulators expect it, but it has quietly become the entire job for a lot of practitioners, and that's a problem nobody in the function wants to say out loud. Paper compliance was built for a slower, less automated world. The world asking for your risk function’s input now needs something else: probabilistic insight that turns uncertainty into measurable exposure, decision thresholds, planning choices, and control responses that improve performance.

uncertainty → exposure → thresholds → decisions → treatment actions → performance 


 

Your GRC Platform Is Not Your Analysis Tool

Here's the uncomfortable part. The software your organization pays for every year, the workflow tool tracking your controls and evidence, was designed around qualitative scoring and compliance checklists. That's what it's good at. It was never built to run a Monte Carlo simulation, model a loss distribution, or tell an executive the actual dollar range they're exposed to if a third-party vendor gets breached. When your entire analytical process lives inside a tool optimized for audit trails, your analysis stops evolving, because the tool defines the ceiling of what you can produce.

The fix isn't waiting for your vendor's roadmap to catch up. It's a deliberate split: keep the GRC platform for what it does well, evidence management, audit trails, workflow tracking, and move your actual risk analysis somewhere built for it. Python has become the default environment for this kind of work in serious risk functions, largely because Monte Carlo simulation, loss exceedance curves, and scenario modeling are a few dozen lines of code away rather than a custom module request to a vendor. R fills the same role for practitioners with a statistics background. Risk quantification tools and practices, as shared by Hernan Huwyler (GitHub, Webpage Risk Quantification Tool), give you a structured, defensible way to translate a vague "third-party risk" into a quantified range of probable financial loss, the kind of output that survives contact with a CFO's questions.

This isn't about abandoning your GRC tool. It's about recognizing that in large organizations, the practitioners who own the analytical layer, not just the workflow layer, are the ones getting pulled into strategy conversations. Everyone else is running the ticketing system for compliance. Pick one risk scenario this quarter, model it properly with a real distribution instead of a 1-to-5 score, and bring that single output to a senior stakeholder. That's a smaller lift than it sounds, and it's the single highest-leverage habit you can build this year.

Show Up Before the Decision, Not After

Ask any product team why they treat GRC as a speed bump and you'll hear some version of the same story: compliance shows up after the decision is already made, holding a checklist, asking for evidence of things nobody planned or budgeted for. That's not a training problem or an attitude problem on either side. It's a positioning problem, and it produces exactly the friction everyone complains about, developers who route around policy, directors who stop reading the emails, business units who treat risk sign-off as a formality to survive rather than an input worth listening to.

The way out is showing up earlier with something more useful than a questionnaire. When a product team is evaluating a new AI vendor, the version of you that gets invited back to the next meeting isn't the one who hands over a fourteen-page due diligence form with a two-week turnaround. It's the one who can quantify the exposure in terms the business already understands: expected loss, probability of a material incident inside the contract term, cost of the control that would cut that probability in half. That's a fundamentally different value proposition than "here's what compliance needs from you," and large organizations reward it accordingly. Budget follows the people who speak business outcomes. Compliance vocabulary, on its own, gets filed and ignored.

This shift also happens to align with where technical depth is becoming non-negotiable rather than optional. A significant share of entry- and mid-level GRC work, evidence collection, control testing, third-party questionnaires, policy review, is exactly the kind of structured, repetitive task that automation and AI tooling are already chewing through faster than most practitioners want to admit. What doesn't automate well is judgment: knowing whether a control actually addresses the risk given the specific business context, evaluating whether an AI model's governance framework holds up technically, or challenging an architecture decision before it becomes a liability nobody can unwind later. Building real fluency in cloud security architecture, AI governance frameworks such as ISO/IEC 42001 and the NIST AI Risk Management Framework, and quantitative risk methods isn't a nice-to-have credential anymore. It's the difference between staying in the evidence-chasing cycle indefinitely and getting pulled into product design and vendor selection conversations where the interesting decisions actually happen.

Challenges GRC Teams Aren't Naming Out Loud

Risk appetite statements sit in almost every governance framework, and almost none of them connect to an actual operational limit. A board approves a paragraph saying the organization has "low appetite for cyber risk", and six months later nobody can point to the dollar threshold, the incident count, or the downtime figure that would trigger an escalation under that statement. The fix is converting every appetite statement into a measurable threshold before it goes to the board for approval, not after: if appetite can't be expressed as a number a control owner can be held against, it isn't appetite, it's a mission statement.

Most key risk indicators in use today are backward-looking by design, counting incidents, breaches, or control failures after they've already happened, which makes them lagging measures dressed up as management tools. A risk indicator that only moves once the damage is done isn't managing risk, it's documenting it. The organizations getting ahead of this are building leading indicators instead: patch latency trending upward before it produces a breach, employee turnover in a control function trending upward before it produces a control failure. That shift, from counting incidents to tracking the conditions that produce them, is where GRC teams start earning a seat in operational planning conversations instead of quarterly retrospectives.

Control rationalization is the maintenance job nobody wants and almost nobody schedules. Years of mapping the same control to five different frameworks, SOC 2, ISO 27001, PCI, a regulatory mandate, produces control inventories bloated with near-duplicates, each tested separately, each consuming audit hours that add nothing to actual risk reduction. A single control rationalization exercise, run once a year, mapping every tested control to every framework it satisfies and retiring the redundant copies, routinely cuts testing effort by a third without touching the organization's actual risk posture. The savings fund the analytical work that keeps getting deprioritized for lack of time.

Risk reporting to top level still leans heavily on narrative slides, three bullet points and a paragraph of prose summarizing key risks this quarter, with no attached number a director could act on independently. A board that reads "cyber risk remains elevated" learns nothing it didn't already assume. A board that reads "expected annual loss from a material cyber incident is estimated between $4M and $22M, driven primarily by third-party access controls" can ask a specific follow-up question and authorize a specific budget. Every recurring board risk report deserves the same test before it goes out: does this sentence give a director something to decide, or something to nod at.

Non-financial compliance and reporting, what most of the market still calls ESG, has become a genuine operational data validation problem disguised as a disclosure problem. Sustainability figures, emissions estimates, supply chain labor metrics, energy consumption by facility, are still frequently collected through spreadsheets emailed between departments with no audit trail, no source-system validation, and no reconciliation against the operational systems that actually generated the underlying activity. Investors and regulators are increasingly treating these numbers with the same scrutiny once reserved for financial statements, which means the same discipline applies: source-system extraction instead of manual entry, a documented chain of custody from raw data to reported figure, and periodic sampling to confirm the numbers tie back to something real. A non-financial disclosure that can't survive an audit trail request is a liability wearing the shape of a report.

Post-incident reviews routinely produce a document, a set of lessons learned, a list of remediation actions, and then that document goes into a folder that nobody connects back to the risk register the incident should have informed. The control gap that caused the incident sat in the risk register the whole time, usually rated lower than its actual severity, and the post-mortem rarely triggers a formal re-rating. Building a hard requirement that every closed incident review updates the corresponding risk entry, with the new severity and likelihood justified by what actually happened, turns incident response from a one-time cleanup exercise into a continuously improving risk model.

Vendor contracts covering AI models, cloud infrastructure, and critical software increasingly include warranty and SLA language that reads as protective but measures almost nothing. "Commercially reasonable security measures" and "industry-standard uptime" are phrases that survive legal review and mean nothing operationally, because neither is tied to a number, a testing cadence, or a remedy that triggers automatically. Contracts covering AI-specific risk should specify measurable performance thresholds, model accuracy bounds, bias testing frequency, incident notification windows measured in hours, not "promptly", with financial remedies that activate without a renegotiation. A warranty nobody can invoke isn't risk transfer, it's the appearance of risk transfer.

Compliance training gets measured almost universally by completion rate, the percentage of employees who clicked through the module, rather than by whether anyone retained anything useful from it. A 98% completion rate on phishing awareness training tells you nothing about whether your organization's actual phishing click-through rate improved, and in most companies nobody bothers to check the second number against the first. Pairing every mandatory training program with a follow-up behavioral metric, simulated phishing results, policy violation rates, control testing outcomes, tracked for the following quarter, is the only way to know whether the training changed behavior or just satisfied an audit requirement.

GRC platform procurement decisions get made by committees evaluating feature checklists against a request for proposal, and the resulting tool frequently sits unused for its most expensive capabilities within eighteen months, because the workflows it assumes don't match how the organization actually operates. The fix isn't a better procurement process, it's sequencing: pilot the platform's core workflow against one real business unit's actual risk process for a full quarter before signing an enterprise-wide contract, and kill the deal if the pilot doesn't produce adoption without heavy-handed mandates. Shelfware is not a technology failure. It's a procurement process that never tested the thing it was buying against real behavior.

The line between first-line risk ownership and second-line risk oversight has blurred badly in organizations that expanded GRC scope faster than they clarified accountability, leaving business units assuming compliance owns their risk and compliance assuming the business unit does. That ambiguity surfaces at the worst possible moment, during an incident, when everyone is asking who was supposed to be watching this. A documented, tool-enforced RACI at the control level, not the department level, specifying exactly which named role owns the risk decision and which role independently assures it, closes that gap before an incident forces the conversation. Clarity here isn't bureaucracy. It's the difference between a fast, coordinated incident response and a room full of people discovering, in real time, that nobody was actually watching the thing that just failed.

The Job Nobody Should Be Doing Alone

There's a pattern that shows up constantly in large companies and gets treated as normal when it absolutely shouldn't be: one person owning policy, NIST compliance, PCI, enterprise risk, security awareness training, third-party risk, IT resilience, incident response, and crisis communications, solo, for an organization with real scale. That's not a lean GRC function. That's an unsustainable accumulation of accountability with none of the authority or resourcing to match it, and it's worth naming as a career risk, not just an operational headache, because when something eventually fails in an environment stretched that thin, and something always eventually does, the person holding every thread is the person holding all the blame.

Organizations that handle this well don't solve it by finding a hero who can do it all. They solve it with structure: a formal RACI that makes ownership and accountability explicit instead of assumed, documented resourcing requirements tied directly to specific control objectives rather than vague headcount asks, and a real risk-acceptance process that requires a senior leader's signature when staffing falls below what the scope actually demands. If you're the one person carrying all of this today and you can't get additional resourcing approved, the professional move isn't to quietly absorb the gap and hope nothing breaks. It's documenting, in writing, that management has accepted the risk of inadequate coverage, and keeping that documentation current as the scope shifts. That single habit protects you individually, and it's also the mechanism that eventually forces the resourcing conversation your organization has been avoiding.

Decision theory already settled this question. A model is worth only what it changes in the outcome of a decision, never how precise its number looks. That is the metric I now apply to my own work. I do not judge a risk model by its confidence interval. I judge it by how many plans, limits, and controls actually move when the distribution moves.

A distribution that never changes a decision is not analysis. It is documentation with better math.

None of this is about doing more with less indefinitely. It's about recognizing where the value in GRC work is actually moving, toward quantification, toward earlier engagement in decisions, toward technical depth that survives automation, and building your practice deliberately in that direction before the gap between paper compliance and real risk advisory becomes the thing that defines the rest of your career.

AI Isn't Coming for GRC Jobs. It's Coming For The Manual Review Part of Every GRC Job

 

What second line experts each need to build before their function gets automated out from under them

Here's the uncomfortable part nobody says out loud in a GRC conference room. AI is not replacing risk managers, compliance officers, auditors, cyber teams, controllers, or sustainability experts. It's replacing the manual review work that used to justify half of those job descriptions. What's left after that work disappears is judgment, and judgment is either your biggest career asset right now or the skill you never actually built because the manual work always came first.

Every one of these six professions is being pulled through the same transformation at the same time, just wearing different clothes. Risk teams are using AI to process larger volumes of exposure data faster than any analyst could by hand. Audit is automating the routine testing that used to eat most of fieldwork season. Cyber teams are automating alert triage and first-line response. Compliance is watching AI surface policy conflicts across thousands of documents in the time it used to take to review one contract. Controllers are automating reconciliations and close procedures. Sustainability teams are automating ESG data extraction and disclosure drafting.

None of that is a headcount story on its own. It becomes one for the people who don't adapt, because the professionals who can validate AI outputs, challenge exceptions, and decide where a human still has to sign off are becoming the only ones a board actually needs in the room.


 

Manual Review Jobs Are Becoming Judgment and Governance Jobs

Guidance on responsible AI and audit puts this plainly: building strong governance, inventories, and validation practices now means an organization can answer the hard questions with confidence when auditors ask them, with a clear account of how AI is actually being used across the enterprise. That's not a compliance platitude. It's a description of what the job becomes once the underlying manual task gets automated: you stop being the person who does the review, and you become the person who can prove the review was done correctly.

The same shift shows up in new studies on audit committees, which stresses that internal auditors still need to bring human judgment into evaluating AI outcomes for fairness, accuracy, reliability, and consistency. The AI does the first pass. The professional's value moves entirely into catching what the first pass got wrong, and knowing when to trust it versus when to escalate.

The riskiest moment in any AI-enabled function isn't when the AI makes a mistake. It's the meeting where everyone assumes someone else already checked the output.

This pattern holds across every one of the six roles, and it's worth naming what "judgment and governance" actually means in practice, because it's not a soft skill. It's validating outputs against known failure modes, challenging exceptions instead of rubber-stamping them, and deciding, explicitly and in writing, where a human still has to own the final call. Professionals who can do all three become harder to automate than the task they used to perform, because the task was never the actual value. The check was.

Redesign The Workflow, Don't Just Bolt on a Tool

The biggest mistake organizations make right now is layering an AI tool onto an unchanged process and calling it transformation. It isn't. Research on operational risk modernization is direct about this: an AI-driven framework is only as good as the data foundation underneath it, and the real opportunity is rethinking the framework itself, not just automating today's manual steps inside the old one.

That distinction matters for your career, not just your organization's efficiency numbers. If you only learn to operate a new tool inside an old workflow, you've picked up a skill that gets replaced the next time a better tool ships. If you learn to redesign the workflow itself, meaning new decision points, new escalation paths, and new control ownership, you've picked up a skill that survives every tool upgrade after this one.

Process design and control mapping are not adjacent skills to model literacy anymore. They're the load-bearing skill. Anyone can learn to prompt a tool. Far fewer people can look at a redesigned workflow and correctly identify where the old control broke, where a new one needs to exist, and who now owns it. That's the professional a board actually wants advising them, and it's a skill you build by practicing control mapping deliberately, not by waiting for it to show up as a side effect of using AI tools daily.

Every GRC Function Now Needs Its Own AI-Specific Controls

Generic "AI governance" is not a control. It's a slogan. Each of the six functions needs controls tuned to its own specific failure modes, because the way AI breaks a compliance workflow is not the way it breaks a cybersecurity workflow.

Compliance officers need to track policy and regulatory drift as a distinct, monitored risk category, not a once-a-year policy refresh. Governance Intelligence's roundup of 2026 GRC predictions captures why this matters: Diligent's governance lead expects the pace of AI regulation to stay unpredictable and increasingly demanding through the year, which means a compliance program built around annual policy review cycles is already structurally too slow for how fast the underlying regulatory landscape is moving.

Auditors need to test AI-enabled controls and the reliability of AI-generated evidence directly, not just the outputs those controls used to produce manually. Internal audit guidance frames this as a genuine fork in the road: internal audit can either lead on AI governance or scramble to catch up after a model failure, compliance breach, or public misstep has already happened. Testing evidence quality now means asking how a model was developed, deployed, validated, and monitored, not just whether the final number tied out.

Cybersecurity experts need controls built for AI-accelerated attacks and AI-driven defense at the same time, because both sides of that fight are now running on the same underlying technology. New cybersecurity surveys name this directly as a defining contradiction facing security leaders: AI is accelerating the threat landscape while simultaneously becoming a core defense capability, which creates pressure to govern adoption tightly without slowing the business down. Establishing a formal AI security and governance program with real human-in-the-loop controls for critical decisions isn't optional anymore. 

Financial controllers need to watch specifically for automation errors bleeding into reporting and approval chains, a risk made sharper by the fact that no binding regulatory standard currently governs AI use in financial reporting audits. Coverage of the 2026 compliance landscape for CFOs and audit committees is blunt about this gap: there's no PCAOB or SEC standard governing AI in audits as of mid-2026, which means the burden falls entirely on the controller's own internal governance to answer questions regulators haven't formally asked yet, questions like which reporting processes use AI, how those outputs get validated, and who signed off on the tools in the first place.

Sustainability experts need to treat AI's effect on ESG data quality and reporting integrity as a governance risk in its own right, not a side benefit of faster reporting. Legal and sustainability coverage of 2026 ESG trends notes that leading teams are already adopting agentic AI to manage compliance work and automate structured data tagging for digital filings, which introduces new governance risk that needs board-level oversight, particularly around whether AI-calculated figures such as carbon footprints or supplier risk scores can actually withstand a regulatory audit. Academic research on ESG disclosure adds a sharper warning underneath that: AI can genuinely improve consistency and scale in sustainability reporting, but it can just as easily formalize and speed up existing greenwashing and disclosure inconsistency if nobody is checking the outputs against source data. An ESG report drafted faster by AI is not automatically a more accurate ESG report. Speed and accuracy are different axes, and AI only reliably improves one of them without deliberate human validation on the other.

Data Quality And Governance Are Now Foundational Career Skills

Every one of these functions runs into the same wall eventually: AI performance is entirely dependent on the data underneath it, and weak data governance creates downstream risk across risk management, compliance, and control functions alike. Research on AI-enabled risk management states this almost as a warning label: without good data, AI is just artificial noise, and clear data governance is the foundation of any effective AI-enabled risk program.

That's not an abstract point. Researchers broader work on rebuilding data governance for the AI era describes a real structural problem showing up across organizations right now: legacy governance models built for structured, static data are struggling under the weight of unstructured inputs, AI-generated outputs, and metadata that shifts constantly, which slows adoption and quietly erodes trust in the outputs everyone's relying on.

The career implication is straightforward, even if it's not the sexy part of the AI story. Professionals who can actually improve data lineage, define clear ownership, and set real monitoring standards are becoming more valuable than professionals who only consume AI outputs and take them at face value. Data stewardship used to be a back-office function nobody wanted. It's becoming a leadership qualification, because nobody can trust an AI-generated risk score, audit finding, or ESG figure without someone accountable for the data quality underneath it.

The Career Premium Goes To Domain Experts Who Can Also Supervise AI

Here's where this gets specific and useful, rather than another generic "upskill in AI" pep talk. The premium isn't going to AI generalists. It's going to domain experts, meaning people who already understand risk, compliance, audit, cyber, financial controls, or sustainability deeply, who then add AI oversight capability on top of that existing expertise.

In cybersecurity, this shift already has names attached to it. Coverage of how agentic AI is reshaping security teams describes the classic Level 1 SOC analyst role turning into an AI supervisor role, where the human reviews agent output, tunes agent guardrails, and focuses on the nuanced investigations the agent stack can't handle on its own. Specialized roles like AI governance specialists, focused on regulatory compliance and internal audit of the AI systems themselves, and AI red teamers, focused on finding flaws through adversarial testing, are becoming distinct career tracks rather than side responsibilities bolted onto an existing security job.

Audit and compliance are moving in the same direction, just with different labels. The routine testing and checklist work is what gets automated first. What's left, and what's growing in value, is higher-order advisory and assurance work: helping the organization decide what AI governance should actually look like, not just confirming a checklist got completed.

Role-Specific Lens: What Each Function Should Actually Prioritize

Risk managers should focus on predictive risk models, response control agents, emerging exposures, faster scenario detection, and real-time monitoring rather than static, backward-looking risk registers. The shift from periodic review to continuous, risk-based monitoring is exactly what model risk research describes as the direction traditional frameworks need to move, since AI systems drift constantly rather than occasionally.

Compliance officers should focus on regulatory mapping, policy drift detection, and exception governance, treating regulatory change as a continuous input rather than an annual refresh cycle. Given how unpredictable AI-specific regulation is expected to stay, a compliance function that only reviews policy once a year is already behind by definition.

Auditors should focus on AI-enabled control testing, evidence reliability, and moving up into higher-value assurance work rather than routine transaction testing. The chief audit executive conversation happening right now, according to new audit committee guidance, is explicitly about how the internal audit function's talent strategy and skill sets need to evolve alongside the technology itself.

Cybersecurity experts should focus on AI-assisted defense, automated triage, and building resilient human oversight into every critical decision point, rather than trying to out-manual an attack surface that's now partly automated on the attacker's side too. Guidance on security management is explicit that human-in-the-loop controls for critical decisions are not optional in a mature AI security program.

Financial controllers should focus on automated close accuracy, reporting integrity, and approval chain controls, given that no binding standard yet tells them exactly how to govern this. That absence of a formal rulebook is not permission to wait. It's the reason controllers need to build their own internal governance now, ahead of whatever standard eventually arrives.

Sustainability experts should focus on data quality, ESG process integrity, and AI governance specifically inside reporting workflows, since the value of faster ESG reporting evaporates the moment a regulator or auditor finds a figure that can't be traced back to a reliable source.

Moves That Actually Build Career Resilience Across GRC Roles


➤ Redesign roles around judgment, not task completion.
 

AI will keep absorbing routine review work. The durable skill is deciding what still needs a human, not doing every task yourself.

Build AI fluency by function, not generically. 

A generic AI training session teaches nobody anything they can use Monday morning. Risk managers, auditors, controllers, and sustainability teams each need role-specific use cases, controls, and prompts tied to their actual daily workflow.

Make control mapping a core, practiced skill. 

Every AI use case should map to a specific control, approval, piece of evidence, and named owner before it scales past a pilot. Professionals who can translate an AI use case into control language directly reduce operational and regulatory risk, which makes them structurally harder to replace.

Treat data quality as career capital, not a back-office chore. 

The people who can improve lineage, completeness, and governance are becoming the ones organizations can't function without, precisely because AI performance depends entirely on the data feeding it.

Learn to manage mixed human-AI workflows deliberately. 

The future isn't full automation. It's people validating, overriding, and coaching AI systems continuously. Knowing when to trust an output, when to challenge it, and how to document that decision is a skill you have to practice, not one that shows up automatically from using a tool daily.

Build a real specialization moat. 

Broad generalists are easier to automate than specialists who combine deep domain expertise with genuine AI oversight capability. Niches like AI risk, AI audit, AI governance, cyber AI defense, and AI-enabled ESG assurance are where the strongest career positions are opening up right now.

Track task-level exposure, not just headcount. 

If a role is losing routine tasks faster than it's gaining higher-value ones, that's a leading indicator, and leaders need to redeploy people into analysis, control, or advisory work before displacement turns into layoffs nobody saw coming.

Become the translator between the business and the AI. 

The professionals who can explain AI risk, AI value, and AI's actual limits in language executives and regulators understand are the ones who become genuinely difficult to replace, because that translation work connects a technology decision directly to accountability.

Push for internal AI champions inside every function. 

A local champion who tests tools, shares lessons, and surfaces risks quickly helps a team adopt AI safely, and gives individual employees a real path to grow into new responsibilities instead of getting overtaken by change decided somewhere else.

Tie AI adoption directly to workforce strategy. 

AI should never be treated as a separate technology program running alongside the workforce plan. Linking AI investment to reskilling, career paths, and role redesign from the start is what lets a workforce evolve with the tools instead of getting overtaken by them.

What This Actually Means for Your Next Twelve Months

None of this requires you to become a data scientist. It requires you to get specific about the same five questions in every AI-touched process you own: what is the metric, what is the threshold, who owns it, how often is it tested, and what happens when it's breached. If you can't answer all five for a control you claim to have, you don't have a control yet. You have a policy statement waiting to fail its first real test.

Start with one workflow you already own. Map where AI has entered it, name the control that used to catch problems there, and check whether that control actually survived the redesign or quietly disappeared along with the manual task it used to sit inside. That single exercise, repeated across a career instead of done once for a compliance checkbox, is the actual difference between a GRC professional AI displaces and one it makes indispensable.

If you're working through this shift in your own function and want to compare notes on what a real AI-specific control catalog looks like for your specific role, that's exactly the conversation worth having now, before the next audit cycle forces it. Subscribe below for the next piece in this series, where we build out the control catalog for each of these six functions line by line.

Probability Is Not Intuition, A Quantitative Risk Framework Every Risk Manager Must Own

 

Why Most Risk Models Break Before the Stress Test Even Starts

A risk manager approved a scenario analysis The model showed a 3% probability of simultaneous credit default and operational system failure. The number felt conservative. The model was wrong. The analyst had multiplied two standalone probabilities together without checking whether the events were independent. They were not. The actual joint probability was nearly four times higher.

This is not an exotic failure. It happens in credit committees, insurance pricing teams, and capital adequacy reviews every week. The underlying error is always the same: treating probability concepts as interchangeable when they are structurally distinct.

The most expensive probability errors in risk management are not computational. They are conceptual. Using an unconditional probability where a conditional one is required, or assuming independence without testing it, can produce capital estimates that understate tail risk by multiples, not percentages.




 


Discrete versus Continuous Random Variables

Before you build a loss model, you need to decide what kind of random variable you are modeling. This choice determines which tools you can use and which results are mathematically valid.

A discrete random variable takes a countable number of values. The number of counterparty defaults in a quarter, the number of operational incidents in a month, and the credit rating of a bond (AAA, AA, A, BBB) are all discrete. You can assign a specific probability to each possible outcome, and those probabilities must sum to exactly one.

Formally, if a discrete random variable X can take values x₁, x₂, ..., xₙ with associated probabilities p₁, p₂, ..., pₙ, then:

P[X = xᵢ] = pᵢ, and Σpᵢ = 1

A continuous random variable can take any value within a range. Annual equity index returns, time to recovery after a system failure, and loss severity on a defaulted loan are continuous. The key consequence: the probability of any single exact value is zero. You cannot ask "what is the probability the loss is exactly $10,432,817?" The answer is always zero. You can only ask about intervals.

The table below captures the practical distinction risk managers need to carry into model selection.

DimensionDiscrete Random VariableContinuous Random Variable
Values it takesCountable, finite or infinite listAny value in an interval
Probability of one exact valueCan be positiveAlways zero
Probability toolProbability mass functionProbability density function
Risk examplesDefault count, claim count, rating categoryLoss severity, time-to-default, VaR level
Sum or integral constraintProbabilities sum to 1Density integrates to 1

Confusing variable type leads to model misspecification. Fitting a continuous distribution to a discrete count variable, or treating a severity measure as discrete, produces biased probability estimates. The decision point is simple: can the variable take non-integer values in principle? If yes, treat it as continuous.


Probability Density Functions: Shape Is Information

For a continuous random variable, the probability density function (PDF) describes the relative likelihood of outcomes across the range of the variable. The PDF itself does not give probabilities directly. Probabilities come from areas under the curve over intervals.

Formally, for a random variable X with density function f(x), the probability of X falling between r₁ and r₂ is:

P[r₁ < X < r₂] = ∫f(x)dx, evaluated from r₁ to r₂

The density function must satisfy two conditions. It cannot be negative at any point. And it must integrate to one across the full range, because something must happen.

A zero-coupon bond example makes this concrete. Define f(x) = x/50 for 0 < x < 10, where x is the bond price. The probability that the price lands between $8 and $9 is:

∫(x/50)dx from 8 to 9 = [x²/100] from 8 to 9 = 81/100 − 64/100 = 17%

The shape of f(x) carries information about where outcomes cluster. A PDF that is steep and narrow signals low uncertainty. A PDF that is flat and wide signals high uncertainty. A PDF with a heavy right tail signals the possibility of extreme positive outcomes. A PDF with a heavy left tail signals the possibility of extreme losses.

When reviewing a loss model, do not focus only on the mean or the single reported percentile. Ask for the full PDF shape. A loss distribution with a thin tail and a fat tail produce identical means but radically different capital requirements. The shape is the risk.


Cumulative Distribution Functions

The cumulative distribution function (CDF) is the workhorse of applied risk quantification. It gives the probability that a random variable is less than or equal to a specific value. Formally:

F(a) = ∫f(x)dx from the lower bound to a = P[X ≤ a]

Three properties of the CDF are worth holding clearly:

The CDF starts at zero at the minimum of the distribution and reaches one at the maximum. It is non-decreasing everywhere. And the derivative of the CDF is the PDF, so you can recover density information from a cumulative function by differentiation.

To find the probability that a variable falls between two values a and b (with b > a), you subtract CDFs:

P[a < X < b] = F(b) − F(a)

To find the probability that a variable exceeds a value a:

P[X > a] = 1 − F(a)

Using the same bond price example, the CDF is F(a) = a²/100. The probability the price lands between $8 and $9 is F(9) − F(8) = 81/100 − 64/100 = 17%, confirming the PDF result through a different calculation path. Both methods must produce identical answers. If they do not, the model has an error.

The CDF is what you use to answer "what is the probability we breach our limit?" or "what is the probability losses stay below our capital buffer?" It is the direct link between a probability model and an operational risk threshold. Build the habit of translating every risk question into a CDF question before running numbers.


Inverse Cumulative Distribution Functions From Probability to Threshold

The inverse CDF runs the calculation backward. Instead of asking "what is the probability of staying below value a?", you ask "what value corresponds to a given probability level p?"

Formally, if F(a) = p, then F⁻¹(p) = a, where 0 ≤ p ≤ 1.

From the bond example, F(a) = a²/100, so solving for a gives F⁻¹(p) = 10√p. At p = 25%, the value is 10√0.25 = 5. Twenty-five percent of the distribution falls at or below a price of $5.

Risk managers encounter the inverse CDF constantly, often without using that name. Value at Risk (VaR) at the 99th percentile is the inverse CDF of the loss distribution evaluated at 0.99. Stress test loss thresholds set at a given confidence level are inverse CDF outputs. Capital adequacy standards that require losses to be covered at the 99.9th percentile require the inverse CDF evaluated at 0.999.

When a model outputs a VaR number or a capital threshold, that number is an inverse CDF value. Understanding this matters because it means the number is only as reliable as the distribution assumption behind it. If the tails of the distribution are misspecified, the inverse CDF at extreme quantiles is wrong, often dramatically wrong. Heavy-tailed distributions produce far larger inverse CDF values at the 99th percentile than normal distributions with the same mean and variance.


Mutually Exclusive Events

Two events are mutually exclusive if they cannot occur simultaneously. A bond cannot be upgraded and downgraded at the same time. A single trade cannot settle and fail on the same date. A counterparty cannot be in default and current at the same moment.

For mutually exclusive events A and B, the probability that either occurs is:

P[A ∪ B] = P[A] + P[B]

This extends to any number of mutually exclusive events: the probability that any one of n mutually exclusive events occurs is the sum of their individual probabilities.

For example, if the probability of a stock return below −10% is 14% and the probability of a return above +10% is 17%, and these two events cannot happen simultaneously, then the probability that the return is either below −10% or above +10% is 14% + 17% = 31%.

The addition rule for mutually exclusive events is simple but easy to misapply. The confusion arises because the English word "or" can mean either "at least one of" (inclusive or) or "exactly one of" (exclusive or), and the formulas differ. In scenario analysis, confirm that your scenarios are genuinely mutually exclusive before summing their probabilities. Scenarios defined by different macro states (recession, stagnation, expansion) are mutually exclusive only if they are exhaustive and non-overlapping by construction.


Independent Events, When Multiplication Is Valid

Two random variables are independent if the outcome of one does not affect the probability of the other. If stock market returns and weather outcomes are independent, then:

P[rain and market up] = P[rain] × P[market up]

This multiplication rule holds only when independence is genuine. A 20% probability of rain and a 40% probability of stock XYZ returning more than 5%, with the two events confirmed independent, gives a joint probability of 20% × 40% = 8%.

Independence and mutual exclusivity are not related concepts. In fact, if both events have nonzero probability, they cannot be simultaneously independent and mutually exclusive. Mutual exclusivity forces the joint probability to zero. Independence, when both events have positive probability, forces the joint probability to be positive. The two conditions are logically incompatible for non-trivial events.

Independence is an assumption, not a default condition. Two credit exposures in the same sector are not independent. Two operational risks sharing the same control environment are not independent. Two market positions driven by the same macro factor are not independent. The most common source of model underestimation in portfolio risk is assuming independence between exposures that are actually correlated. Validate independence assumptions against historical joint outcomes before relying on simple multiplication.


Joint Probability and Probability Matrices

Joint probability is the probability that two events occur together. For independent events, the joint probability is the product of the marginal probabilities. For dependent events, it requires more information about the relationship between the two variables.

A probability matrix organizes joint probabilities in a table where rows represent outcomes of one variable and columns represent outcomes of another. Each cell contains the joint probability of the row outcome and column outcome occurring together. Row and column totals give the marginal (unconditional) probabilities of each variable separately. All cells must sum to one.

A bonds-and-stock example demonstrates the mechanics. Consider a company with bonds (upgrade, no change, downgrade) and equity (outperform, underperform). The joint probability of bonds being upgraded and stock outperforming is 15%. The marginal probability of stock outperforming, found by summing down the outperform column, is 50%.

When cells in the matrix are missing, they can be recovered using the row and column total constraints. If the outperform column must sum to 50% and already shows 5% and 40%, the missing cell is 5%. That recovered value can then be checked by confirming the row total equals the known row marginal.

Probability matrices are underused in enterprise risk management. A matrix crossing credit states (upgrade, stable, downgrade) against market regimes (bull, neutral, bear) gives immediate visibility into whether risks are concentrated in dangerous joint states. A cell showing a 12% joint probability of "corporate downgrade" and "market stress" is far more actionable than two separate 30% probabilities reported in isolation.


Conditional Probability: Updating Risk Estimates with New Information

Conditional probability is the probability of event A given that event B has already occurred. The formula is:

P[A | B] = P[A ∩ B] / P[B], provided P[B] > 0

The vertical bar means "given." P[market up | rain] reads as "the probability the market is up, given that it is raining."

Conditional probability and joint probability are connected through this formula. Rearranging gives:

P[A and B] = P[A | B] × P[B]

This is equally valid written as:

P[A and B] = P[B | A] × P[A]

Both forms are mathematically equivalent. Which form is more useful depends on what information you have and what you are trying to estimate. This distinction becomes central in Bayesian analysis, where you update probabilities as new information arrives.

The link between conditional and unconditional probability runs through the law of total probability. If a random variable X can take values x₁ through xₙ, then the unconditional probability of any event Y is:

P[Y] = Σ P[Y | xᵢ] × P[xᵢ]

In words: the overall probability of Y is the weighted average of the conditional probabilities of Y given each possible state, weighted by the probability of each state.

Conditional independence is a related concept. If the probability of the market being up on a rainy day equals the probability of the market being up on a dry day, then the market is conditionally independent of rain. Formally:

P[market up | rain] = P[market up | no rain] = P[market up]

When conditional independence holds, the joint probability of two events equals the product of their marginal probabilities. When it does not hold, multiplication produces the wrong answer.

Conditional probability is what separates reactive risk management from predictive risk management. The unconditional probability of a counterparty default may be 2%. The conditional probability of default given a two-notch rating downgrade in the prior 90 days may be 18%. Those two numbers require entirely different responses. Monitoring, escalation, and hedging decisions should be driven by conditional probabilities, not unconditional ones.


A Concrete Risk Management Example: Credit and Operational Risk Combined

A regional bank's risk team is reviewing whether to include operational risk and credit risk in a combined stress scenario. The standalone probability of a significant credit loss event (defined as losses exceeding the 95th percentile of the credit loss distribution) is 5%. The standalone probability of a major operational failure event is 3%.

The team initially models the joint probability as 5% × 3% = 0.15%, assuming independence. The capital calculation rests on that number.

A closer review finds that both risks share a common driver: a core banking system outage. When the system fails, credit monitoring controls are also impaired, which elevates default detection latency. The events are not independent.

Using a joint probability matrix built from 10 years of incident history, the team finds the actual joint probability of simultaneous credit loss and operational failure events is 0.9%, six times the independence-based estimate.

The capital implication is material. The tail loss in the joint scenario requires additional buffer allocation. The original model, built on an untested independence assumption, would have left the bank undercapitalized for a scenario that history shows is not negligible.

The corrective step requires no exotic mathematics. It requires correct use of a probability matrix, a test of the independence assumption against historical joint frequencies, and the conditional probability framework to update estimates when a leading indicator (system degradation signal) is observed.


The Practical Decision Framework for Risk Probability Questions

Every quantitative risk question maps to one of five probability tools. Knowing which tool answers which question eliminates most conceptual errors before they reach a model.

Risk QuestionCorrect ToolWhat to Watch For
What is the shape of our loss distribution?Probability density function (PDF)Tail shape, skewness, multimodality
What is the probability we breach a limit?Cumulative distribution function (CDF)Distribution assumption in the tails
What loss corresponds to a target confidence level?Inverse CDFTail sensitivity to distribution choice
What is the probability two risks occur together?Joint probability / probability matrixIndependence assumption validity
What is the probability of loss given a trigger?Conditional probabilityConditioning event definition and data quality

Governance risk: The most common audit finding in quantitative risk models is not a computational error. It is an undocumented assumption. Independence assumptions, distribution choices, and conditioning events should be explicitly stated, tested against historical data, and reviewed when the economic environment changes. An assumption valid in a low-correlation regime can fail catastrophically in a stress regime.


What Risk Managers Should Do with This Framework

Start by auditing your current portfolio models for independence assumptions. Identify every place where joint probabilities are computed as products of marginals. For each one, ask whether historical co-occurrence data supports the independence assumption. Flag any case where a shared macro driver, shared control environment, or shared counterparty makes independence implausible.

Build probability matrices for your top five combined risk scenarios. Put credit states on one axis and operational or market states on the other. Populate the cells from historical frequency data, not from assumed independence. The matrix will immediately show you where joint risk is concentrated.

Switch your escalation triggers from unconditional probabilities to conditional ones. If a counterparty's credit spread widens by 150 basis points, the relevant number for your response is not the unconditional default probability. It is the conditional default probability given that spread move. That conditional probability should drive your monitoring intensity, hedge sizing, and reporting escalation.

Finally, when reviewing any risk model that outputs a quantile-based metric (VaR, Expected Shortfall, capital at risk), ask two questions: what distribution assumption drives the inverse CDF? And has that assumption been back-tested at the tail, not just at the center of the distribution? Most model risk in quantitative finance lives in the tails, exactly where the inverse CDF is most sensitive to distributional choice.




Subscribe for More Quantitative Risk Frameworks

This publication covers the mathematical and statistical foundations of risk management, the governance structures that make quantitative models reliable, and the operational failures that happen when probability theory is applied carelessly at scale.

The next articles in this series cover covariance and correlation in portfolio risk, Bayesian updating for early-warning systems, and the specific failure modes of normal distribution assumptions in fat-tailed loss environments.

If you are building, reviewing, or governing quantitative risk models, subscribe now. The technical depth here is written for risk managers who need to understand the machinery, not just the outputs.

S/4HANA Role Redesign: Fix Segregation of Duties Before Go-Live or Pay the Audit Bill After

 

Why Privilege Creep Kills SAP Migrations Before the First Production Transaction Runs

Article by Prof. Hernan Huwyler, MBA, CPA, CAIO
AI GRC Director | AI Risk Manager | Quantitative Risk Lead
Speaker, Corporate Trainer and Executive Advisor

Top 10 Responsible AI and Risk Management by Thinkers360Your migration to SAP S/4HANA is six months out. The project team is focused on data migration, Fiori tile configuration, and cutover planning. Meanwhile, 847 user roles built across eight years of organizational changes, job transfers, emergency firefighter access, and M&A integrations are being lifted wholesale into the new system. Nobody has reviewed them. Nobody has mapped them against the new S/4HANA authorization model. And your external auditors are already asking for the SoD conflict report.

This is the standard failure mode. And it is expensive to fix after go-live.

Migrating unremediated ECC roles into S/4HANA production does not just inherit old access risk. It amplifies it. S/4HANA's simplified data model, new Fiori authorization objects, and transaction replacements create net-new SoD conflicts from role content that was previously clean.

This article gives you the technical remediation workflow to stop that from happening. It covers the SAP-native tools, the sequencing logic, the role design architecture that prevents re-accumulation, and the automated tooling that makes the process viable at enterprise scale.


Why S/4HANA Breaks Your Existing SoD Matrix

SAP S/4HANA fundamentally changes the authorization landscape. This is not an incremental upgrade. The SAP S/4HANA Simplification List documents thousands of transaction replacements, program removals, and process consolidations. Transactions you built roles around in ECC no longer exist, have merged into Business Partner (BP) transactions, or now trigger different authorization objects entirely.

The most visible examples: XD01 and XK01, the customer and vendor master maintenance transactions, are replaced by the SAP Business Partner transaction BP. Any role that controlled access through those old transactions now either fails to control the equivalent S/4HANA function or, worse, grants broader access than intended because the authorization check structure has changed.

SAP Community analysis of S/4HANA security landscape changes confirms that SoD matrices built for ECC become unreliable after conversion. The conflict rules reference transactions that no longer exist, miss new authorization objects introduced in S/4HANA, and fail to account for Fiori app authorizations that bypass the traditional transaction-based access model.

This is not a configuration problem. It is a design-time engineering problem.


What Privilege Creep Actually Looks Like at Migration Time

Privilege creep is the accumulation of access rights that users no longer need. It happens across three vectors in most enterprise SAP environments.

First, job transitions. When a user moves from accounts payable to controlling, they get new access. The old access rarely gets removed. After five years and three job changes, that user holds access spanning procurement, finance, and HR, a combination that violates SoD in ways no single access request ever triggered.

Second, project-based access. Implementation projects, year-end processes, and audit support cycles generate temporary elevated access. Firefighter IDs get reused. Temporary roles become permanent. Emergency access granted during a system outage in 2021 is still active in 2025.

Third, role sprawl. When role designers copy existing roles rather than build from a business capability model, each copy carries forward every transaction and authorization value from the original. Over years, roles accumulate dormant transactions that nobody uses but that still represent active access rights from a SoD perspective.

By the time an S/4HANA migration project starts, a mid-sized SAP environment typically carries hundreds of roles where 30 to 50 percent of assigned transactions have zero usage in the prior 12 months. 

Migrating roles with unused transactions into S/4HANA does not just carry old risk forward. In some cases, S/4HANA's new authorization objects cause those dormant role contents to map to broader access than they did in ECC. A transaction that was harmless in ECC may activate a wider authorization check in S/4HANA.


The SAP-Native Remediation Workflow: SU24, SUIM, and PFCG in Sequence

You do not need a third-party tool to start. SAP provides three native transactions that, used in the right sequence, give you a complete picture of your current access exposure before you touch a single production role.

Step 1: Use SUIM to Inventory Current Access

The SAP User Information System, transaction SUIM, is your starting point for understanding who has what access across the landscape. SUIM lets you query users by role, by authorization object, by transaction, and by combinations of those dimensions.

Run four reports at minimum before remediation begins. Pull all users with access to posting transactions (FB01, VF01, MIGO) combined with approval transactions in the same process area. Pull all users with access to sensitive BASIS transactions such as SU01, SU10, and SE16N. Pull all roles that include more than 150 active transactions, which is a strong indicator of over-provisioned design. And pull all user-role assignments where the last logon date is more than 90 days ago, which surfaces dormant accounts that should be disabled before migration.

SUIM does not tell you whether combinations are SoD violations. It tells you the raw access landscape. That data feeds your SoD analysis.

Step 2: Identify SoD Conflicts Against a Current Ruleset

Your SoD ruleset must reflect S/4HANA transactions, not ECC transactions. If you are using a GRC Access Control ruleset that was last updated during your ECC 6.0 implementation, it will miss conflicts introduced by S/4HANA's process changes. Update the ruleset first, then run the conflict analysis against the SUIM output.

SAP GRC Access Control provides the standard enterprise framework for this analysis. It maintains business process rule libraries, maps conflicting function pairs, and generates conflict reports by user, role, and profile. For organizations without GRC, the same analysis can be run manually using SUIM data cross-referenced against a documented SoD matrix, but at enterprise scale that is not a sustainable approach.

Research from Gutesman et al. on real-time SoD conflict detection establishes the theoretical basis for pre-runtime SoD checking: identifying violations before they are committed to production is significantly less costly than detecting and remediating them after user provisioning is complete. The same principle applies to migration projects.

Step 3: Correct Authorization Defaults in SU24

Before you rebuild roles in PFCG, fix the authorization defaults that PFCG uses as its source of truth.

Transaction SU24 maintains the default authorization objects and their proposed values for each transaction code and each Fiori application. When a role designer adds a transaction to a role in PFCG, PFCG reads SU24 to know which authorization objects to include and what default values to propose.

If SU24 defaults are incorrect, every role built from them inherits the error. In S/4HANA migrations, SU24 defaults for replaced or new transactions may not reflect your security requirements. Review SU24 for every transaction in scope, particularly for Fiori apps that were not present in your ECC system. Set proposal indicators correctly. Mark authorization objects that should always be checked. Remove defaults for authorization objects that your policy excludes.

This step is frequently skipped in migration projects under time pressure. That decision consistently produces roles where authorization checks are either missing or over-permissive, because PFCG built them from uncorrected defaults.

Step 4: Rebuild Roles in PFCG from Business Capabilities

Do not copy roles from ECC. Build them from business function definitions.

SAP's role maintenance transaction PFCG is where roles are constructed, maintained, and generated. A role built correctly in PFCG contains only the transactions and authorization objects required for a defined business function, with org-level values appropriate to the user population.

Use a three-layer architecture. Single roles contain the smallest functional unit, such as "Create Purchase Order." Composite roles bundle single roles into job profiles, such as "Procurement Clerk for Plant 1000." Derived roles replicate the design of a master role across different organizational units without duplicating the permission structure.

This architecture limits role count, simplifies audit reporting, and makes future modifications predictable. When a business function changes, you modify one single role, and the change propagates to every composite and derived role that references it.

Step 5: Validate and Test Before Cutover

After rebuilding roles, run the full SoD conflict analysis again against the new role set. Confirm that all high-priority conflicts identified in Step 2 are resolved. Then run authorization trace analysis, using SAP transaction ST01 or the authorization check framework documented in SAP's authorization evaluation documentation, to confirm that users can complete their required business processes without hitting authorization failures.

SAPinsider's analysis of go-live security sequencing makes the cost case clearly: remediating SoD conflicts before cutover costs a fraction of the effort required after production users are live and dependent on incorrect access. Post-go-live remediation also introduces operational risk, because correcting over-permissive access in production can break business processes that users have already built workarounds around.


When to Start the Role Work: A Sequencing Framework

The timing question is not primarily a technical decision. It is a resource and risk decision.

Before the S/4HANA Project Starts

For organizations with large, complex role landscapes, start role remediation before the S/4HANA project formally begins. This phase focuses on ECC cleanup: removing unused transactions, resolving existing SoD conflicts, standardizing role naming and structure, and establishing the governance model that will govern the S/4HANA design.

Starting early means the S/4HANA project inherits a cleaner baseline. It also means the project team can focus on S/4HANA-specific changes rather than simultaneously debugging both legacy role problems and new S/4HANA authorization requirements.

The main cost argument against early start is that some role work done in ECC will need to be redone when S/4HANA-specific transaction changes are applied. That is true. It is still cheaper than the alternative.

During the S/4HANA Project

Once the S/4HANA development system is available, apply the migration-specific changes. Map ECC transactions to their S/4HANA equivalents using the Simplification List. Add Fiori app authorizations. Test role content in the S/4HANA environment, because authorization behavior can differ from ECC even for transactions that exist in both systems.

This phase should address only S/4HANA-specific changes if the pre-project cleanup was done correctly. If it was not, this phase will be significantly more complex and will compete for time with every other workstream in the project.

After Go-Live

Post-go-live role work should cover edge cases, fine-tuning based on actual user behavior, and deferred items that were explicitly scoped out. It should not be the primary remediation phase. Organizations that defer SoD remediation until after go-live consistently face audit findings within the first annual review cycle.


The Non-Human Identity Problem in SAP Authorization

Role design discussions in SAP environments almost exclusively focus on human users. This is increasingly the wrong frame.

Modern SAP landscapes run significant automated workloads: interface users for middleware connections, batch users for background job execution, RFC users for system-to-system communication, and service accounts for cloud integration platforms. These non-human identities frequently hold broad authorizations granted during implementation and never reviewed since.

Research by Poreddy on non-human identity governance frameworks establishes that traditional role-based access control and HR-driven lifecycle models are structurally inadequate for managing machine identities. These accounts do not have managers who receive access review requests. They do not appear in HR offboarding workflows. They accumulate permissions across system upgrades without anyone noticing.

In S/4HANA migration projects, interface and batch accounts are typically migrated without review because the migration team is focused on human user profiles. Those accounts then exist in production with ECC-era authorizations that may map to broader S/4HANA access than was intended.

A batch user account with S_TABU_DIS access to table maintenance in ECC may, in S/4HANA, gain access to new configuration tables that did not exist in the source system. The authorization object and its values are identical. The scope of access has expanded.

The remediation approach for non-human identities follows the same SUIM-SU24-PFCG sequence used for human users, with one addition: every non-human identity should have a documented owner, a defined technical purpose, and a review cycle independent of HR processes. That governance structure is absent in most SAP environments today.


Automated Role Design: Where Authorization Architect Fits

At enterprise scale, manual role redesign is not a viable path. A large SAP environment with 2,000 or more roles, multiple system landscapes, and a six-month migration timeline needs tooling that automates the mechanical steps while maintaining human control over design decisions.

Transaction Usage Analysis as the Design Evidence Base

Authorization Architect uses Transaction Archive tool to analyze actual SAP execution history across the landscape. The recommended observation window is 13 months, long enough to capture a full annual business cycle plus overlap for month-to-month variation. This matters because role cleanup based on shorter windows risks removing access that is genuinely needed but only exercised during annual processes such as year-end close or tax reporting.

The analysis outputs a fact-based view of which transactions in each role are actually used, which are dormant, and which appear only in role definitions but have never generated an authorization check in production. This replaces assumption-based role cleanup with evidence-based redesign.

Performance: the Transaction Archive analysis typically completes in under two minutes for standard role sets, often in seconds. Present this as an operational benchmark for interactive review sessions, not as a contracted service-level commitment.


Authorization architect user dashboard

SoD Screening Before Role Creation

Authorization Architect runs SoD analysis against the proposed role design before the role is built, not after it is assigned to users. This is the architecturally correct sequence. Detecting a conflict in a role proposal costs nothing. Detecting the same conflict in a production role assigned to 200 users costs weeks of remediation effort.

The SoD check uses a functionality against either the client's existing ruleset or the default library. It covers sensitive transaction checks as well as function-pair conflicts. The result is visible to role designers before they submit the role for approval.

Basin et al.'s foundational work on dynamic enforcement of separation of duty constraints established that workflow-independent SoD policy enforcement, applied at design time rather than runtime, provides the most effective compliance architecture. Authorization Architect's pre-creation SoD check applies this principle directly to SAP role engineering.

Workflow Approvals Embedded in the Role Build Process

Authorization Architect routes role proposals through a structured approval workflow. Role operators create proposals and submit them for review. Role owners, the business stakeholders responsible for role content, receive pre-populated approval requests showing the proposed transactions, org levels, and SoD analysis results. Role approvers authorize composite roles for production deployment.

This three-tier structure, operators, owners, approvers, enforces the separation of design and approval duties that most SAP security governance frameworks require but few organizations actually enforce technically.

Automated Role Build Beyond SU24 Recommendations

When a role proposal clears the approval workflow, Authorization Architect builds the role automatically. This includes naming and description, role long text, structured role menu, org level definitions, and the full authorization object set. It generates master and derived role pairs and creates standalone maintain, display, and composite job roles as appropriate for the design.

The automation goes beyond what SU24 recommends. It applies the organization's naming standards, incorporates the corrected SU24 defaults from the remediation workflow, and produces consistent role content regardless of which team member runs the build. That consistency is critical for audit purposes and for future maintenance.

Migration Provisioning for S/4HANA Transaction Replacements

For migrations specifically, Authorization Architect's Migration Provisioning feature maps ECC transactions in existing roles to their S/4HANA equivalents using the Simplification List. Where a transaction no longer exists in S/4HANA, the tool suggests the replacement. Where a transaction remains, it flags it for testing confirmation.

This is the systematic approach to the XD01-to-BP and XK01-to-BP problem described earlier. Instead of manually tracing each transaction through the Simplification List, the tool processes the full role inventory and produces a remediation plan showing every impacted transaction and its proposed resolution.


Authorization Governance After Go-Live: Sustaining What You Built

Role redesign before go-live solves the migration problem. It does not solve the ongoing governance problem.

Access accumulates again after go-live. New projects generate emergency access. Organizational changes produce new provisioning requests. Integrations add non-human identities without formal review. Within 18 to 24 months of a well-executed migration, many organizations are back to a state of meaningful privilege creep if they have not built the governance infrastructure to prevent it.

Bhatia's research on AI-driven compliance architecture in S/4HANA transformations identifies continuous surveillance and dynamic regulatory data feeds as the structural components that prevent post-go-live compliance drift. The point is sound: compliance-by-design at migration time creates the clean baseline, but sustaining it requires automated monitoring that catches new violations before they accumulate into systemic risk.

The practical controls to build into your post-go-live governance model are direct:

Connect provisioning workflows to HR triggers. Every hire, job change, and termination should automatically generate an access review or modification request. Access that does not get reviewed on a triggered basis will not get reviewed.

Run quarterly access reviews for critical roles and high-risk users, not annual reviews. Annual reviews are too slow to catch the accumulation patterns that create material audit exposure.

Use transaction usage data from the production system, the same data that Authorization Architect analyzes at design time, as your ongoing evidence base for access review decisions. Roles with significant unused transaction populations after six months of production operation should be flagged for cleanup.

Set automated deprovisioning rules for dormant accounts. Accounts that have not logged in for 90 days should be flagged. Accounts inactive for 180 days should be disabled, pending review.

Treat non-human identity governance as a separate program, not a subset of human user access management. Document every interface user, batch user, and RFC user with an owner, a technical purpose, and a review schedule. Review those accounts on the same quarterly cycle used for privileged human users.


A Concrete Example

Procurement SoD Remediation Before Cutover

A North American manufacturing company with 4,200 SAP users and 1,800 roles began its S/4HANA migration with a SUIM-based access inventory. The analysis identified 23 users in the procurement organization who held simultaneous access to purchase requisition creation (ME51N), purchase order approval (ME29N), and goods receipt posting (MIGO). All three functions in one user profile is a textbook SoD violation covering the full procure-to-pay cycle without any compensating control.

Further analysis showed that 14 of the 23 users had accumulated this combination over time through separate access requests, each individually approved, none evaluated against the combined access picture. The remaining 9 held roles that had been copied from a prior template without SoD review.

The remediation sequence: SUIM identified the population. SoD analysis confirmed the conflict severity. Role redesign in PFCG split the three functions into separate single roles, assigned to different user populations with clear functional separation. Authorization Architect ran SoD screening against each redesigned role before build. The three composite roles were approved through the workflow process and built automatically.

Total elapsed time from analysis to production-ready roles: four weeks. The same remediation started post-go-live, with 4,200 live users dependent on their current access, would have required business sign-off on access removal for active users, parallel testing to prevent process disruption, and a change management effort to explain why users were losing access they had held for years. The post-go-live version of the same work typically takes three to five months.


Key Failure Modes to Avoid

Copying ECC roles into S/4HANA without transaction mapping. Copied roles carry every legacy transaction, including those that no longer exist or have changed authorization behavior. This produces both access gaps (missing S/4HANA functions) and access excess (dormant ECC transactions that now trigger different authorization checks).

Running SoD analysis after user provisioning, not before role creation. Post-provisioning SoD analysis identifies conflicts but requires user-level remediation, which is operationally disruptive and business-politically difficult. Pre-creation SoD analysis stops conflicts from entering the design.

Treating non-human identities as out of scope for migration role review. Interface, batch, and RFC users carry real access risk. Their authorizations should be reviewed against the S/4HANA authorization model with the same rigor applied to human user roles.

Deferring SoD remediation to post-go-live. The cost differential between pre-go-live and post-go-live remediation is not linear. Post-go-live remediation competes with operational support, requires business process testing to avoid disruption, and generates audit findings in the first review cycle if not completed quickly.

Using Authorization Architect as a role-copy factory. Automated role build accelerates execution. It does not replace design judgment. Automation run against an uncorrected business function model or a flawed SoD ruleset produces wrong roles faster.


What to Do Next

If your S/4HANA project is in planning, start the SUIM inventory now. Pull the four reports described in Step 1 of the remediation workflow. That data will tell you the scope of your remediation problem before the project schedule is locked and the role work gets compressed into the final two months before cutover.

If your project is already underway, the same analysis applies with more urgency. Get SUIM data, run SoD conflict analysis against a current ruleset, and prioritize the highest-severity conflicts for resolution before the transport to production.

If you have already gone live and the role work was deferred, build the quarterly review cadence and the non-human identity inventory program now. The clean-up will take longer than pre-go-live remediation would have. Starting immediately limits the accumulation.

The technical workflow is not complicated. The governance discipline required to execute it consistently, across teams, landscapes, and project timelines, is where most organizations struggle. That discipline starts with a clear decision that SoD is a design-time engineering responsibility, not a post-go-live audit finding.


Subscribe to Continue This Work

This publication covers SAP authorization engineering, identity governance, and the technical architecture of secure enterprise systems. Future articles will go deeper on Fiori authorization design, SAP GRC Access Control ruleset maintenance, and non-human identity governance frameworks for cloud-integrated SAP landscapes.

If you are building or fixing authorization governance in an SAP environment, subscribe now so these articles reach you directly. No generalist technology news, no vendor marketing. Just the technical depth that the people actually doing this work need.