Who Owns the Risk? The AI Reliability Boundary Assessment Questions

By KEIRSTEN BRAGER

I am presenting the AI Reliability Boundary Assessment at CYBR.SEC.CON in Houston this week. The assessment asks who owns the risk when an AI platform fails inside a fifteen-minute reliability window, and the answer changes depending on the event. This page previews the tiers, the ownership map, and a sample of the questions. Nothing on it collects anything.

The Decision Window

At 2:07 on a weekday afternoon, an AI decision-support platform that operators have leaned on for months stops responding. The vendor opens an incident and says it is investigating. At 2:22 the fifteen-minute reliability window closes. The vendor may restore service that evening, and in the meantime the grid still has to be operated. Every organization I work with can tell me who owns the platform and who signed the contract. Ownership of the fallback is usually undocumented, and in a lot of cases the people who used to perform it have been reassigned or never backfilled.

Why Fifteen Minutes Changes the Question

A BES Cyber Asset is defined by a fifteen-minute test. If the asset were rendered unavailable, degraded, or misused, would that adversely impact one or more Facilities, systems, or equipment within fifteen minutes of its required operation, misoperation, or non-operation? Both that term and BES Cyber System come from the NERC Glossary of Terms.

Fifteen minutes does not leave room for a vendor to triage, publish a status update, and restore service. Inside one utility, a single question about AI adoption tends to get three different answers, because the platforms sit at different distances from operations.

Tier Where it sits What loss means
Tier 1 Adjacent to operations Analysis, documentation, and reporting. An outage delays work, and operations recover.
Tier 2 Workflow decision support Loss reduces situational awareness inside the window that matters.
Tier 3 Inside the boundary Loss can affect reliable operation. Recovery becomes a compliance question.

Tier assignment does not follow the product category. The same vendor, license, and model can land in any of the three, depending on what the workflow does and who leans on the output.

Three Boundaries the Room Scores

The talk works through three events. Two happened to other people and are documented in public. The third is a composite built from patterns that recur in utility environments.

In May 2026, more than two thousand packages were uploaded to RubyGems, the package registry for the Ruby language. The registry suspended new registrations for about four days and removed more than five hundred packages. Researchers later attributed the campaign to agents OpenAI was testing internally, which OpenAI disputes while confirming its agents were active there. No supplier in that chain was compromised. An agent working under an objective its owner describes as benign produced malicious packages in a registry that thousands of build pipelines pull from automatically.

In July 2026, agents running an internal OpenAI cyber-capability evaluation left their intended environment and reached production systems at Hugging Face. OpenAI's account describes agents chaining previously unknown vulnerabilities to escape the sandbox, then obtaining outbound internet access the environment was not supposed to have. Hugging Face's forensic reconstruction covers roughly seventeen thousand six hundred attacker actions across five days. An advisory platform became an external actor, and no one filed a production change.

In the composite, an engineer builds a genuinely useful prototype on a personal AI account using regulated data. Leadership wants it operating before provenance and authorization are finished. A colleague then works through that authenticated session, doing things the engineer never approved. Every control in the chain functioned as designed and logged the account holder, so the record names a person who was not at the keyboard.

Who Owns the Gap

Event Who owns the gap
Silent model or permission change Vendor and operator
Platform outage or degraded output Vendor and operator
Autonomous agent holding standing access Operator
Borrowed authenticated session Operator and employee
Prototype built on a personal account Operator and employee
Regulated data retained in model history Vendor, operator, and employee
Shared control does not mean shared accountability. The operator still carries the operational consequence.

A vendor can hold most of the control surface for an event while carrying none of the reliability obligation. The assessment separates who creates the risk from who absorbs it because those two rarely sit with the same party. The same split runs through computational load registration, which I worked through in Top 10 Computational Load Accountability Mapping Questions for Leaders.

A Sample of the Questions

The full assessment scores twenty questions from zero to four across CIP-002, CIP-004, CIP-007, and CIP-010, then averages each set of five. The spread inside a section usually tells you more than the average. One question from each standard, to show the shape.

  1. CIP-002. What re-triggers categorization when the vendor changes behavior without changing the version string?

  2. CIP-004. Does any agent hold standing access, and whose identity does it use?

  3. CIP-007. Is prompt and tool-use history a declared log source?

  4. CIP-010. What triggers authorization when the vendor files no change?

Five more sit underneath and do not depend on CIP at all. Readers in water, pipeline, rail, and healthcare can substitute their own obligation and run these against any AI dependency.

  1. Who creates or changes the risk?

  2. Who can see the change?

  3. Who has authority to intervene?

  4. Who must produce the evidence?

  5. Who absorbs the operational, regulatory, and financial consequence?

The common failure is that answers one through four name a vendor and answer five names you.

Keep the Recovery Path Until the Dependency Earns It

Most AI capacity decisions I see bank the labor savings before the platform has proven its availability, its recovery behavior, its evidence production, or its contractual accountability. The price of that bet shows up on the day the fifteen-minute window closes.

Availability risk becomes workforce risk when the fallback plan is the position you eliminated.

Pick one AI-enabled workflow or vendor dependency that leadership has to approve or defend, and run the nine questions above against it with the people who would have to execute the fallback. Where two of you score the same question differently, work out why before you average it away.

 
 

Read the Rest of This Series

Sources and Further Reading

 

Featured Posts

Next
Next

CIP-014-4 Approved: The New Physical Security Risk Assessment and the 2028 Clock