K · Practical guide

Practical Guide: Energy and Utilities

Two estates with different physics: an operational technology side where errors are physical and cascading, and a customer side where errors take someone's heating away.

Typical top tier
Tier 4 — anything touching switching, dispatch, protection or disconnection
Already-owned ground
Functional safety, OT security zoning, control-room procedure, incident reporting
Hardest gate
G2 — because the architecture decision is also a safety-zone decision
First record to fix
The authority boundary for anything that can issue an operational instruction

1. Where the risk actually concentrates

Utilities run two AI estates that share a logo and nothing else. On the operational side, outputs become physical actions on plant and network: switching, curtailment, dispatch, protection settings, maintenance intervention. Errors here are irreversible in the strongest sense the framework defines, they propagate through a network rather than stopping at a customer, and the existing governance is functional safety and OT security rather than model risk.

On the customer side, outputs determine debt treatment, disconnection, prepayment, vulnerability classification and tariff. The consequence of an error is not a physical hazard but a household without power or heat, often a household already known to be vulnerable. The two estates need the same framework and quite different anchors.

The overlay's most important rule is architectural rather than procedural: an AI system must not be a safety function. Protection, interlocks and safety-instrumented systems stay deterministic and stay independent of the model. AI may inform, prioritize, predict and propose. Where it acts, the acting path must sit behind the same independent protection that a human operator does.

Forecast load, generation, price Tier 3 Operate switching, dispatch, curtailment Tier 4 Protect settings, interlocks — no AI in path Tier 4 Maintain prediction, intervention order Tier 3 Serve billing, debt, disconnection Tier 4
Two estates, one framework. The operational columns are governed by physics and by functional safety practice; the customer column is governed by consumer protection. Both reach Tier 4, for unrelated reasons. The protect column is the one place where the correct architecture decision is to keep AI out of the path entirely.

Indicative classification of the systems this sector keeps building. The scores are illustrative, not authoritative: they show how the anchors in Chapter 6 read against sector facts. Score your own system; do not copy a row.

Typical system Illustrative D1–D5 Tier What is usually mis-scored
Automated switching, reconfiguration or DER dispatchD1 4 · D2 3–4 · D3 4 · D4 4 · D5 2Tier 4D5 read as low because the data is telemetry; uncertainty about behaviour is the risk, not sensitivity
Protection setting or interlock recommendationD1 4 · D2 2 · D3 4 · D4 4 · D5 2Tier 4Advisory scoring in a control room where the recommendation is accepted under alarm load
Wildfire, flood or asset-failure risk driving a public safety shutoffD1 4 · D2 2–3 · D3 4 · D4 4 · D5 3Tier 4Both error directions are severe, and only one of them is measured
Predictive maintenance and intervention prioritizationD1 3–4 · D2 2–3 · D3 3 · D4 3 · D5 2Tier 3A deferred intervention is a decision, not an absence of one
Load, generation and price forecasting into tradingD1 3 · D2 3–4 · D3 3–4 · D4 3 · D5 2Tier 3–4Positions taken on a forecast are financially irreversible once the market clears
Outage prediction and crew dispatchD1 3 · D2 3 · D3 2–3 · D4 4 · D5 2Tier 3Restoration order is a decision about who stays without power longest
Debt, disconnection and prepayment decisionsD1 4 · D2 2 · D3 4 · D4 4 · D5 4Tier 4Treated as a collections process; a wrong disconnection in winter is a safety event
Meter anomaly and theft detectionD1 3–4 · D2 2 · D3 3 · D4 4 · D5 3Tier 3An accusation of theft is close to irreversible for the customer relationship
Consumption analytics and tariff recommendationD1 2–3 · D2 2 · D3 2 · D4 4 · D5 3Tier 2–3Half-hourly consumption data is re-identifying and reveals occupancy

2. The regulatory interface

Regulatory note. The regimes below are named so each IRGF record can be pointed at the obligation it evidences, not to restate them. Applicability, thresholds and commencement dates differ by jurisdiction and several have moved during implementation. Nothing here is legal advice: confirm the current position with your own counsel, and record the answer in the Regulatory Overlay Reference so it is checkable later.

IRGF does not restate any of these obligations. It gives each one a record that carries the evidence, an owner, and a trigger that reopens it when the obligation or the system changes.

Regime or standard What it obliges in practice IRGF record that carries the evidence
Critical infrastructure protection and OT security regimes (for example NERC CIP in North America, NIS2 in the EU, and sector licence conditions)Asset identification, electronic security perimeters, change management, access control, incident reporting, and supply chain risk management for systems affecting reliable operation.Zone and conduit placement recorded in the AI-extended ADR; incidents routed through the existing OT incident process, not a parallel AI one
Industrial security and functional safety standards (IEC 62443, IEC 61508 / 61511)Zone separation, safety integrity levels, independence of protection layers, and management of change.The architectural rule that AI is never a safety function, recorded as a pattern constraint in the pattern library
Wholesale market integrity rules (for example REMIT in the EU, market manipulation prohibitions elsewhere)Prohibition on manipulative strategies, inside information handling, and records of trading decisions.Strategy constraints in the agent authority boundary; decision records retained to the market standard
Consumer protection and supply obligations, including protections for vulnerable customersRestrictions on disconnection, obligations to identify and support vulnerable households, and complaints handling.Vulnerability handling recorded as a D1 justification and a control, with the detection path evidenced
Smart metering and data protection rulesPurpose limitation on consumption data, consent for granular readings in some regimes, and security requirements.Data Lineage and Sensitivity Record for every meter data flow, with the granularity recorded
EU AI Act, where it appliesAI as a safety component in the management and operation of critical infrastructure appears in the high-risk list.Regulatory Overlay Reference, mapped per operational system

3. Calibrating the five dimensions

The dimensions do not change. What changes is what a 3 and a 4 look like when the subject matter is this sector, and which reading an assessor under delivery pressure reaches for first.

Dimension How to read it here The mis-score to watch for
D1 Decision ConsequenceAnything that can de-energize, energize, curtail, disconnect or delay restoration is 4. On the customer side, loss of supply to a household with a known vulnerability is 4.Splitting the estate's anchors by department rather than by consequence, so the customer side scores itself against commercial impact.
D2 AutonomyControl-room acceptance under alarm load is not review. Closed-loop remediation is 4 regardless of the operator's ability to intervene afterwards.Scoring the designed intervention point rather than the one available during an event, which is when the system matters.
D3 Reversibility DeficitPhysical actions, market positions and disconnections are 4. Equipment damage and safety consequences do not reverse because the switch can be closed again.Mechanism reversibility again: the breaker reclosing is not the same as the outage not having happened.
D4 Exposure and ScaleNetwork topology is the scale factor. A single bad instruction propagates along the network, and cascade potential should push D4 to 4 for anything on the transmission or primary distribution path.Counting connected customers rather than network reach and interdependency.
D5 Sensitivity and UncertaintyOperational telemetry is low sensitivity but often high uncertainty. Half-hourly consumption data is sensitivity 3. Take the higher sub-score.Scoring OT systems as D5 = 1 because no personal data is involved, and losing the uncertainty half of the dimension entirely.

4. One system, end to end

The overlay above is a map. This is one route across it: a single fault followed from detection to post-event review, with the record or control that attaches at each step.

Worked example — automated network reconfiguration after a fault D1 4 · D2 4 · D3 4 · D4 4 · D5 2 → IMPACT 4 · CONTROL DEFICIT 4 → TIER 4 IN THE WORLD IN THE RECORD, AND WHAT WATCHES IT Fault detected from operational telemetry Staleness limit per feed: a model on stale telemetry fails closed to the deterministic path Model proposes a switching sequence Zone and conduit placement recorded in the ADR; protection stays independent of the model Instruction executes inside the approved envelope Agent Card: instruction scope, rate limits, and what the agent may never issue Network state changes; supply is restored or lost Kill-switch pre-approved at G3, with an out-of-band stop that does not run through the loop Event reviewed with the control room Drift watch: boundary divergence, instruction volume, operator acceptance under alarm load
Where the framework says no. Column two carries the one architectural prohibition this sector should never trade away: protection and interlocks stay deterministic and independent. Everything to its right assumes that holds — a kill-switch that depends on the automation being healthy is not a kill-switch.

5. What each gate adds

Additions only. Everything in the base gate definitions still applies; see the gate checklists for the common set.

Gate Sector addition Why it is here
G1State which side of the OT boundary the system sits on and whether any output can become an operational instruction. Confirm no safety function is being delegated.The answer determines the security regime, the review path and the tier floor, and it is expensive to change later.
G2Zone and conduit placement, protection independence, and the failure mode when the model is unavailable. Every operational data source needs a staleness limit, because stale telemetry is worse than none.Availability and staleness are safety-relevant in a way they are not in most sectors.
G3Control-room procedure updated and rehearsed, kill-switch conditions defined for closed-loop systems, and an agreed maximum instruction rate.Authorization has to include the human procedure. The system and the control room are one system.
G4Any change to autonomy, instruction scope or rate limits is Material. Seasonal reconfiguration counts as change.Operational configuration drifts continuously, and it is the parameter that determines consequence.
G5Retention aligned to asset and incident investigation timescales, which are long, and to regulatory record-keeping for market activity.Investigations into physical events reach back years.

6. Controls and evidence worth adding

Control Where it attaches Evidence it produces
Architectural prohibition: no AI in a safety function or protection pathPattern library constraint; ADR conformance checkPattern conformance recorded at G2, with any deviation escalated rather than self-certified
Instruction rate and scope limits for any acting systemAgent Card, enforced at the control layerMachine-readable boundary comparable against actual instruction logs
Telemetry staleness monitoring with a defined maximum ageData Lineage Record; runtime signalAge-of-data metric per feed, with the action on breach stated
Operator acceptance-rate monitoring in the control roomRuntime signal; assurance cycleAcceptance rate by shift and by alarm load, which is where it degrades
Vulnerability check before any disconnection or supply restrictionControl Matrix; workflowEvidence the check ran and what it returned, per decision
Both-direction error monitoring for shutoff and curtailment decisionsAssurance cycleMeasured false-positive and false-negative rates, reported together

7. Runtime signals to wire first

Control Plane onboarding order matters more than coverage in the first year (Chapter 18). These are the signals that earn their place earliest in this sector.

Signal Drift category Suggested response
An AI service gains a network path into an OT zone it was not approved forSecurity / integrationImmediate: this is a perimeter breach in the terms the OT regime already uses
Operator acceptance rate rises during high-alarm periodsBehavioral, with a classification consequenceRe-score D2 for the event condition, not the average condition
Telemetry feed exceeds its staleness limitData lineage and governanceFail closed to the deterministic path; a model on stale data is worse than no model
Instruction volume or scope exceeds the approved envelopeAgent boundary divergenceIncident at any tier: suspend the acting path, preserve state, notify the control-room manager
Disconnection or debt-action rate changes by customer segmentBehavioralRoute to the customer vulnerability lead and the regulator-facing team together

8. Failure modes this sector produces

Advisory in the control room, autonomous in the storm

The tell. During normal operation the operator considers each recommendation. During a major event, with three hundred alarms an hour, every recommendation is accepted.

The response. Score and instrument the event condition separately. If the system will be trusted absolutely during events, govern it as an acting system and put the protection independence in place accordingly.

The pilot that reached into the process network

The tell. A cloud analytics service is given a read path into the historian, and six months later a write-back capability is enabled to close the loop.

The response. Zone placement is an architecture decision recorded at G2. Enabling write-back is an autonomy change and a Major change, not a configuration flag.

Forecast confidence taken as certainty

The tell. Point forecasts are consumed by downstream optimization that has no representation of uncertainty, and positions or dispatch decisions are taken as if the forecast were fact.

The response. Require the uncertainty interval to cross the interface, and treat its absence as an assurance failure rather than a modelling nicety.

Collections automation meeting a vulnerable household

The tell. A debt model is efficient, accurate on its own terms, and disconnects a customer whose vulnerability flag was in another system.

The response. Make the vulnerability check a preventive control in the action path, not a report reviewed monthly.

9. A ninety-day start

If the sector is yours and the framework is new, this is the order that produces something defensible fastest. It assumes one part-time architect and one risk lead, not a programme.

  1. Split the estate into OT-side and customer-side and give each its own scoring session with its own anchors.
  2. List every system whose output can become an operational instruction, directly or through an integration. Include anything with a write path to a historian or SCADA layer.
  3. Confirm in writing, per system, that no safety function depends on a model. Where one does, that is the first remediation and it outranks all documentation work.
  4. Score the top three operational systems and the debt or disconnection system, with OT engineering and the customer vulnerability lead in the room.
  5. Set staleness limits on the operational feeds that matter and wire the alert.
  6. Write authority boundaries and kill-switch conditions for any closed-loop remediation.
  7. Measure control-room acceptance rate for a month, separating event periods from normal running.
  8. Run a G3 dry run on the highest-tier operational system with the control-room manager present.