The AI Reliability Boundary: A Black Hat Debrief for the Grid

By KEIRSTEN BRAGER

Almost every conversation at Black Hat came back to AI, and almost every pitch assumed the answer to an AI problem is another product. In the grid, that assumption does not hold. Here’s our debrief on what the show floor missed, why governance is not a document, and why AI is not automatically out of scope for NERC's Critical Infrastructure Protection standards. It closes with the questions to answer before your next vendor demo.

Overview

I spent the week at Black Hat, and the trend was impossible to miss. Almost every conversation came back to AI. Vendor after vendor promised to detect it, prevent it, and govern it. A growing number claimed to respond to it on their own. Some of the startups were impressive. I do not want to take anything away from them. The monitoring gaps around AI are real. Vendors are racing to catch up on privilege escalation and anomalous agent behavior. Some of those detection rules still have to be built. That work matters, and it is coming.

AI did not own the whole conference. There were robots, playful animal themes, and strong sessions on infrastructure, agents, and the future of how we work. While the AI security conversation dominated the floor, I left it with a different takeaway than the vendors were selling.

The pitch on the floor assumes the answer to an AI problem is another product. In the grid, that assumption does not always hold. The Federal Energy Regulatory Commission (FERC) has directed the North American Electric Reliability Corporation (NERC) to develop reliability standards for large computational loads, the category that will cover many AI data centers, with the registry criteria and first standards due at FERC by December 31, 2026. If the term Computational Load Entity is new, our plain-English guide breaks it down. AI and grid reliability now sit on the same line. I call that line the AI Reliability Boundary, where AI governance meets grid reliability. Everything I carried home from Las Vegas runs back to it. By the end, ask yourself three things, depending on your seat.

Capacity and Capability

Most utilities are not short on tools. They are running the tools they own at thirty or forty percent of capacity. Meanwhile, plenty of organizations are still working to get baseline controls fully in place. Adding one more console to that picture does not raise your security. It adds another thing to manage on top of a stack nobody is using fully.

Then there is the harder gap, and it is not on any price sheet. Securing and governing AI is a different skill set. It is not a bolt-on to the day job. Management should not assume the existing team can simply absorb this work. Those teams are already staffed lean and carrying everything else they own. It is not fair to them, and it sets them up to fail. In the early phase, this needs dedicated capability, not another list of tasks squeezed between other duties. Build it deliberately across people, process, and technology. Collaborating with the other teams by design, yes. Quietly dumped on people at capacity, no. AI security requires specific expertise, real training, and dedicated focus. That expertise exists to identify, manage, and continuously reduce the gaps.

I said it on the panel and I will say it here.

The hottest technology available to most utilities is the stack they already own. It just has to be matured.

Why Any of This Matters

It helps to remember why we do this work. Security does not exist to serve the security department. We are not here to admire our metrics on vulnerabilities found and closed. We are here so the business can keep delivering reliable electricity. That is the whole job.

The business is not waiting on us. AI is being deployed faster than it is being governed. Often it ships with no governance at all. The priority has been to get it into production and sort out governance later. The result is predictable. Automation ends up running in places nobody inventoried. Workflows carry inherited, long-lived, high-privilege tokens to code repositories and cloud services. AI orchestration spins up a short-lived identity, does its work, and disappears. It leaves stale permissions behind.

Nicole Carignan of Darktrace and I kept landing on the same point. These agents are starting to look like insiders. Not insiders with motive, which is what our programs are built around. These are insiders with access, speed, and a vulnerability to manipulation.

An insider with no intent is still an insider.

Governance is Not a Document

On the panel, Nicole made the case for security before governance. I agree. I have always believed that security done correctly enables compliance by default. Governance is the layer that comes after. It is the mechanism that confirms the controls you built are still working. Which brings me to a word that gets used loosely.

Governance is not a document. Picture an AI usage policy sitting on a file share. No technical control can detect when someone deviates from it. That policy is not weak governance. It is the absence of governance.

If you want to know whether you actually have a governance program, run three tests. Is there a detective control that enforces the policy, or only the policy? Has someone decided and documented who owns the risk when an agent misbehaves? That includes an agent that is over-permissioned, escapes its container, or acts without authorization. Do the people responsible for oversight have any AI security or governance training?

If the answer to those is no, you have an org chart, not a program.

Done right, governance is a running loop, not a one-time sign-off. It means automated detection when something deviates from your AI policy. It means security controls tested on a schedule, not assumed. It means risk assessments on a cadence, confirming the controls still work as designed. A program that does those things is alive, not framed on a wall.

The good news is you already own a model for this. You govern change every day. Some activities are pre-authorized as business as usual. Others must clear a change advisory board. Apply the same thinking to agents. Someone has to define what business as usual means for agentic activity. What can an agent do routinely? What requires authorization? What can never happen without a human in the loop? Critical operational technology (OT) processes demand defined human accountability and clear authorization boundaries, even when automation is involved. Someone has to know how to respond when automation does something unexpected.

Before you buy anything, run a risk assessment against the tools you already have. Find out how much of this they can do today. Can they identify an agent? Can they tell agent activity from human activity? Can they flag a deviation from your AI use policy? A privilege escalation? A spike in traffic to a destination no human chose? Some of your existing toolset already does more of this than you think. Some of it does not, which is worth knowing before you buy more.

A Word on Scope

Some organizations have decided AI is simply out of scope for NERC's Critical Infrastructure Protection (CIP) standards. I would be careful there.

Determining applicability. Do not declare AI out of scope just because it is not a BES Cyber Asset, the Bulk Electric System (BES) term for a device whose loss would affect reliable operation within 15 minutes. Absence from your CIP-011 BES Cyber System Information (BCSI) program is not a reason to treat AI as irrelevant to governance. The real work is determining which CIP boundary and obligations actually apply. Walk that determination deliberately before you conclude anything.

The cloud question. If your AI runs in the cloud, watch the draft CIP 100-series standards taking shape now under NERC Project 2023-09. The first five drafts are posted for informal comment through August 21, 2026, and a first ballot is not expected before 2027. They signal where cloud compliance is heading. They are not an obligation that binds you today.

The work ahead. There is a standard-by-standard way to work through this. It runs from asset inventory under CIP-002, to supply chain under CIP-013, to internal network security monitoring under CIP-015. I will lay it out in a follow-up. For now the point is simpler. Do not assume the boundary does not apply to you.

The Questions We Left With

All of this is really one question wearing different clothes. Each thread lands on the boundary. The boundary is not held by a product. It is held by people with the right skills. It is held by baseline controls applied deliberately to AI. It is held by decisions someone actually made and wrote down.

To be clear, I am not arguing against AI. I am arguing for adopting it securely and governing it properly. Daniel Wallance of McKinsey and Company argued for using AI where automation makes sense. That includes remediating vulnerabilities where possible. I agree. The distinction I would add is for our environments specifically. In OT, a patch carries obligations before it ever reaches production. CIP-007 governs how security patches are evaluated and applied on a set cycle. CIP-010 governs the change itself, authorized and documented against a configuration baseline, and for high impact systems, tested first where that is technically feasible. That discipline does not disappear because an agent is the one applying it. Someone still has to decide and document where automation makes sense. Applying one blanket standard across everything is not a plan.

We left the panel with two questions I do not think anyone has fully answered. What does patching look like when agents talk to agents? The change window and the human operator may not be there anymore. How do we defend against an insider that has no intent? I do not have tidy answers. I have a strong conviction about where to start. Build the capability, not just the tool. Staff the skill set. Adopt AI on purpose. Those are the open questions from the stage. The ones that matter more are the questions you carry back to your own desk.

If you are a CIP Senior Manager:

  • Which AI platforms and agents touch your BES Cyber Systems right now?

  • Is any of your BCSI being processed by a cloud-hosted AI model reachable from the internet? What controls do you have to detect if it is?

  • Who approved that access, and can you prove it?

  • If an agent acted outside its authorization tonight, would you know?

If you allocate capital:

  • Are you funding a new tool, or the capacity to run it?

  • Have you budgeted for the AI skill set, not just the software?

  • Have you priced the choice between a cloud model and a local one?

  • What unowned risk are you carrying that no one has put a number to?

If you lead governance, risk, and compliance:

  • Can you detect a deviation from your AI policy, or only publish it?

  • Are your controls tested on a cadence, or assumed?

  • Have your AI vendors and platforms been through your risk assessment process?

  • Can you attest to where each model came from and what it trained on?

Answer these before the next vendor demo. Holding the boundary starts there.

Follow the series. The next installment walks the boundary standard by standard, from asset inventory through supply chain to internal network monitoring.


Sources and Further Reading

Featured Posts

Next
Next

Your Environment, Your Risk: Building an OT Risk Management Program You Can Defend