Your Environment, Your Risk: Building an OT Risk Management Program You Can Defend
By James Brosnan
There is no rulebook for securing an OT environment, and there never could be. Two utilities can make entirely different choices and both be defensible. Here is how to build an OT risk management program you can stand behind: map your environment, rank the consequences you will not tolerate, choose controls deliberately, and document why, so you can defend every decision an auditor asks about.
Overview
There is no chapter-and-verse rulebook for securing an operational technology environment, and there never could be. A generation operator, a transmission utility, and a distribution cooperative all run on electric-sector OT, but their assets, threats, consequences, and constraints look nothing alike. Even two utilities of the same type (one running a modern EMS, one nursing decades-old substation relays) carry very different risk. That is the first thing to accept about an OT risk management program: no one hands you the answer. You survey your own environment, decide which risks deserve your resources, and live with where those decisions lead. Two organizations can make entirely different choices and both be defensible, because "defensible" is set by their environment, their consequences, and their risk appetite, not by a universal checklist.
The catch is ownership. Those choices are yours, which means you have to make them deliberately, document them, and be ready to explain them to leadership, to business partners, to insurers, to regulators, and to your own incident review when something goes wrong. This is exactly where NERC oversight is headed. Audits increasingly turn on whether you can walk an auditor through your reasoning (not just produce an artifact). A program built to answer "why did you decide this?" is a program built for where the electric sector is already going.
Know the Environment You're In
Everything starts with a current, trustworthy asset inventory, and most OT organizations don't have one. Not really. The single most common gap in industrial environments is knowing what is actually connected. What each device does, what talks to what, what firmware it runs, and, critically, what happens to the physical process if it fails or is compromised.
This is where OT risk management diverges sharply from its IT cousin. In IT, the crown jewels are data and the nightmare is a breach. In OT, the crown jewels are physical processes, and the nightmare is loss of view, loss of control, equipment damage, environmental release, or someone getting hurt. Your inventory has to capture that. Not just IP addresses and operating systems, but process context. A vulnerable engineering workstation matters differently depending on whether it can push settings to protective relays and reconfigure a substation or merely pull reports from the historian.
You cannot make sound decisions about, or defend, an environment you haven't mapped. Visibility comes first for a reason. Every downstream choice, from where to monitor to what to segment, depends on knowing your own network better than an adversary does. Start here, even if it is unglamorous, and accept that the first version will be wrong in places. A rough inventory you correct over time beats a perfect one you never finish.
Define What "Bad" Looks Like
Before you can manage risk, you have to define it, and in OT, risk is best expressed as consequence to the operation. Not CVSS scores, not vulnerability counts, but consequences.
Sit down with system operators, protection engineers, and safety (operations: the people who actually run the grid) and ask the uncomfortable questions. What events would we truly not accept? An unplanned customer outage or forced load shed? Damage to a large power transformer with an 18-month replacement lead time? A protection scheme disabled or misconfigured so a fault doesn't clear? An arc-flash injury to a field crew? Loss of visibility or control at the control center? Rank them. This becomes your consequence framework, and it is the compass for every choice that follows.
This is also where your priorities become genuinely your own. A transmission operator will hold the integrity of protection and control above almost everything, because a mis-operation can cascade. A distribution utility may weigh restoration speed and customer-minutes-interrupted more heavily. A generation operator worries about unit trips and equipment damage. None of them is wrong. What is wrong is never having the conversation so that priorities get set by default, usually by whoever bought the last tool.
Understand the Threats You Face
Threats give your risk assessment its shape. You don't need a classified briefing to do this well. You need honest answers to a few questions. Who would plausibly target us, and why. Ransomware crews looking for downtime leverage? A disgruntled insider with substation or control-center access? A state actor with a demonstrated interest in grid operations? Or simply a commodity worm wandering into a flat network?
Then, what pathways exist into our OT environment? Vendor remote access to relays and RTUs? IT/OT interconnections at the control center? Transient laptops carried between substations? The supply chain behind our field devices?
Then pair threats with your consequence framework. A high-consequence asset, a critical substation or the EMS, reachable through a well-worn pathway by a plausible actor is what deserves your attention. A theoretical vulnerability on an isolated device that cannot affect the grid is noise. A mature program spends its energy on the former.
Supply chain deserves particular attention, because it is the pathway utilities most often underestimate. The relays, RTUs, and software running your grid come from vendors, get serviced by vendors, and are patched by vendors, and every one of those relationships is a way in. Sourcing a component domestically is not the same as being able to trust it; capacity is not assurance. Managing this risk well is not a one-time vendor questionnaire filed away for an audit. It is an ongoing program: knowing who touches your systems, understanding what access they hold, and assessing the risk of the products and services themselves. Treat supply chain as a living part of your risk picture, not a form.
Account for What's Mandatory
Not every input to your program is a free choice. For most electric utilities, NERC CIP is the headline obligation, and its reach depends on how your BES Cyber Systems are categorized: high, medium, or low impact. Beyond CIP you may also answer to FERC orders, state public utility commission requirements, and contractual or insurance terms that carry their own security conditions. These are constraints on the path, not the whole map, but ignoring them has consequences of its own, such as financial penalties (CIP violations are assessed per violation, per day), enforcement actions, lost coverage, or lost contracts.
The shift worth internalizing is that NERC's own oversight has become risk-based. This means your risk program and your compliance program are no longer separate tracks. Inherent Risk Assessments increasingly drive what gets audited and how deeply, Compliance Oversight Plans create traceability between identified risk and monitoring activity, and internal-controls evaluation is now continuous rather than a point-in-time snapshot. The regulator is, in effect, asking you to run the same risk process it runs. Organizations that keep a "risk program" and a "compliance program" that barely speak to each other are duplicating effort and missing the point.
A good program does three things with regulation:
Map requirements to the risks they address. Most mandates exist because a regulator concluded a given risk was serious enough to make a control non-negotiable. Connect each requirement back to the consequence it is meant to prevent and it stops being a box to check and becomes part of your risk narrative. Often a single well-designed control satisfies several obligations at once.
Treat the baseline as the floor, not the ceiling. Meeting a standard has never been the same as being secure, and a requirement rarely maps perfectly to your environment. NERC CIP tells you the minimum the industry agreed on. It cannot know which of your substations keeps you up at night. Where your own assessment says the mandated control is insufficient, you should definitely do more. Where a requirement doesn't reach a risk that matters to you, cover it anyway. Regulators write for the average entity, not for yours.
Document the overlap once. The same asset inventory, consequence framework, and control decisions that drive your risk program are the evidence an auditor wants to see. Build them once, in one place, and let both the risk story and the compliance story draw from the same source. That is what turns audit prep from a fire drill into a report you can run.
Handled this way, regulation stops competing with risk management and starts reinforcing it. You satisfy the people you answer to and you spend the rest of your effort on the risks the rulebook never anticipated.
The CIP Roadmap Is Really a Risk Document
If you want proof that the regulator is thinking this way, read NERC's CIP Roadmap. Released in January 2026, it is not a compliance guide or a retrospective. It is a forward-looking blueprint for how CIP has to evolve, and at its core it is a risk-prioritization exercise. NERC built it on a formal risk registry and scoring model that weighed likelihood, impact, and mitigation maturity across dozens of cyber and physical risk categories, then let that analysis decide where to act first. That is the same method this post describes, run at the scale of North America.
The result is a deliberate move away from scoping controls purely by asset impact tier and toward scoping them by risk and program maturity. The three priorities NERC pulled to the top (multi-factor authentication for remote access, foundational cyber hygiene, and protecting control traffic that rides public or carrier telecom networks) were not chosen because they are novel. They were chosen because the risk math said they buy the most systemic risk reduction for the effort. That is prioritization by consequence.
Two of those three are things a sound risk program is already doing. Foundational cyber hygiene, in NERC's own framing, means asset inventory, network boundary definition, configuration and patch management, identity control, and knowing what is actually connected: the same visibility-first foundation this post opened with. NERC's blunt conclusion is that most of the sector's residual risk is not exotic; it is weak inventories, undefined boundaries, and thin visibility, and advanced monitoring cannot work reliably on an environment nobody understands. If you have done the inventory and consequence work already, you are not bracing for these standards. You are ahead of them.
The Roadmap also widens what counts as your risk. Low-impact systems are no longer treated as low risk, because coordinated attacks across many small assets can now produce system-level effects. And the grid's trust boundary now reaches DER aggregators, inverter-based resources, large controllable loads, cloud control platforms, and vendors with remote access. Parties that create bulk power system risk whether or not they are registered. For a registered entity, that means the supply-chain and third-party exposure discussed earlier is not a side file to be managed later. NERC has signaled it will hold you accountable for the risk those relationships introduce, effectively making you a cybersecurity gatekeeper for your own ecosystem. Pulling those actors into your risk picture now is cheaper than being told to do it later.
The practical read is simple. The direction of regulation has converged on the way this post says to build a program. Entities that already run on a defensible asset inventory, a consequence-based view of risk, and documented control decisions will find that the coming standards (MFA across remote access, telecom-path protection, cloud governance, and hygiene baselines for low-impact systems) are things their program already anticipates. Entities waiting for standards language to tell them what to do will be reacting to a risk picture NERC has already drawn.
Choose Your Controls, and Record Why
Now come the actual decisions. For every significant risk you have the classic four options: mitigate it, transfer it, avoid it, or accept it. All four are legitimate. The immaturity is not in accepting a risk; it is in accepting a risk silently, without analysis, ownership, or an expiration date.
A few principles for choosing well in OT:
Fit the control to the environment, not the other way around. Patching monthly may be trivial in IT and impossible for a substation relay or an EMS you can only touch during a planned outage window. That doesn't mean you shrug. It means you choose compensating controls. Segmentation, monitoring, hardening, and strict access control around the device you cannot patch on IT's schedule.
Prefer controls that protect the grid, not just the network. Protection and control redundancy, independent backup control schemes, and the ability to fall back to manual switching are risk controls too, often better ones than another security appliance. (think Cyber Informed Engineering)
Sequence matters. You cannot monitor a network you haven't mapped, and you cannot segment traffic you don't understand. Work in order: visibility, then segmentation and access control, then detection and response.
Every choice gets an owner, a rationale, and a review date. The one-sentence version of program maturity is “We know what we decided, why, who owns it, and when we'll look at it again.”
That documentation is not bureaucracy. It is the difference between a program and a pile of tools, and it is what lets you defend your decisions later to a board, an insurer, an auditor, or an incident review team.
Plan for Prevention to Fail
A risk management program that only prevents is half a program. Some risk you didn't address (or couldn't afford, or never imagined) will eventually be the one an adversary exploits. Plan for it.
That means detection tuned to what matters in your environment. Unexpected relay setting changes, unexplained breaker operations, new devices on a substation LAN, remote access outside maintenance windows. It means an incident response plan that system operators have actually rehearsed, with clear, pre-decided authority to isolate a substation, fall back to manual switching, or operate the grid without SCADA, settled in advance, not negotiated mid-crisis. And it means recovery you have proven. Tested backups of relay settings, RTU and IED configurations, and EMS/SCADA configuration, plus the procedures to restore them. In OT, resilience (the ability to keep the lights on safely and recover quickly) is often the highest-value risk investment you can make.
Treat the Program as Living
Risk isn't static, and neither is your environment. New interconnections appear, vendors change, threat actors shift tactics, and last year's accepted risk quietly grows teeth. A living program builds in feedback loops: periodic reassessment, metrics that track risk reduction rather than tool deployment, lessons folded back in from incidents and near-misses, and a standing conversation among security, operations, and leadership.
This is also where a risk-based approach outperforms a checklist, and the standards themselves are moving this way. Newer CIP requirements increasingly ask entities to make and defend risk decisions rather than follow prescribed steps, and NERC's oversight has been built to evaluate exactly that kind of reasoning. When the landscape shifts, a risk-driven program can reallocate. It can retire a control that no longer addresses a top risk or redirect budget to the newest one. Checklist-driven programs keep marching down obsolete paths because the checklist says so. And when the auditor arrives, the conversation is no longer "show me the artifact" but "walk me through your reasoning," which is precisely the conversation a real risk program is built to have.
The Bottom Line
Building an OT risk management program isn't about buying the right products or copying another company's playbook. It is about doing the work with discipline. Map your environment, define your intolerable consequences, understand your adversaries, choose your controls deliberately, prepare for prevention to fail, and revise as the world changes. No two organizations will, or should, end up in the same place. The only truly bad outcome is the one you stumble into because you never made real choices at all. The choices are yours. Make them so you can defend them.