Practical Guide: Energy and Utilities
Two estates with different physics: an operational technology side where errors are physical and cascading, and a customer side where errors take someone's heating away.
- Typical top tier
- Tier 4 — anything touching switching, dispatch, protection or disconnection
- Already-owned ground
- Functional safety, OT security zoning, control-room procedure, incident reporting
- Hardest gate
- G2 — because the architecture decision is also a safety-zone decision
- First record to fix
- The authority boundary for anything that can issue an operational instruction
1. Where the risk actually concentrates
Utilities run two AI estates that share a logo and nothing else. On the operational side, outputs become physical actions on plant and network: switching, curtailment, dispatch, protection settings, maintenance intervention. Errors here are irreversible in the strongest sense the framework defines, they propagate through a network rather than stopping at a customer, and the existing governance is functional safety and OT security rather than model risk.
On the customer side, outputs determine debt treatment, disconnection, prepayment, vulnerability classification and tariff. The consequence of an error is not a physical hazard but a household without power or heat, often a household already known to be vulnerable. The two estates need the same framework and quite different anchors.
The overlay's most important rule is architectural rather than procedural: an AI system must not be a safety function. Protection, interlocks and safety-instrumented systems stay deterministic and stay independent of the model. AI may inform, prioritize, predict and propose. Where it acts, the acting path must sit behind the same independent protection that a human operator does.
Indicative classification of the systems this sector keeps building. The scores are illustrative, not authoritative: they show how the anchors in Chapter 6 read against sector facts. Score your own system; do not copy a row.
| Typical system | Illustrative D1–D5 | Tier | What is usually mis-scored |
|---|---|---|---|
| Automated switching, reconfiguration or DER dispatch | D1 4 · D2 3–4 · D3 4 · D4 4 · D5 2 | Tier 4 | D5 read as low because the data is telemetry; uncertainty about behaviour is the risk, not sensitivity |
| Protection setting or interlock recommendation | D1 4 · D2 2 · D3 4 · D4 4 · D5 2 | Tier 4 | Advisory scoring in a control room where the recommendation is accepted under alarm load |
| Wildfire, flood or asset-failure risk driving a public safety shutoff | D1 4 · D2 2–3 · D3 4 · D4 4 · D5 3 | Tier 4 | Both error directions are severe, and only one of them is measured |
| Predictive maintenance and intervention prioritization | D1 3–4 · D2 2–3 · D3 3 · D4 3 · D5 2 | Tier 3 | A deferred intervention is a decision, not an absence of one |
| Load, generation and price forecasting into trading | D1 3 · D2 3–4 · D3 3–4 · D4 3 · D5 2 | Tier 3–4 | Positions taken on a forecast are financially irreversible once the market clears |
| Outage prediction and crew dispatch | D1 3 · D2 3 · D3 2–3 · D4 4 · D5 2 | Tier 3 | Restoration order is a decision about who stays without power longest |
| Debt, disconnection and prepayment decisions | D1 4 · D2 2 · D3 4 · D4 4 · D5 4 | Tier 4 | Treated as a collections process; a wrong disconnection in winter is a safety event |
| Meter anomaly and theft detection | D1 3–4 · D2 2 · D3 3 · D4 4 · D5 3 | Tier 3 | An accusation of theft is close to irreversible for the customer relationship |
| Consumption analytics and tariff recommendation | D1 2–3 · D2 2 · D3 2 · D4 4 · D5 3 | Tier 2–3 | Half-hourly consumption data is re-identifying and reveals occupancy |
2. The regulatory interface
Regulatory note. The regimes below are named so each IRGF record can be pointed at the obligation it evidences, not to restate them. Applicability, thresholds and commencement dates differ by jurisdiction and several have moved during implementation. Nothing here is legal advice: confirm the current position with your own counsel, and record the answer in the Regulatory Overlay Reference so it is checkable later.
IRGF does not restate any of these obligations. It gives each one a record that carries the evidence, an owner, and a trigger that reopens it when the obligation or the system changes.
| Regime or standard | What it obliges in practice | IRGF record that carries the evidence |
|---|---|---|
| Critical infrastructure protection and OT security regimes (for example NERC CIP in North America, NIS2 in the EU, and sector licence conditions) | Asset identification, electronic security perimeters, change management, access control, incident reporting, and supply chain risk management for systems affecting reliable operation. | Zone and conduit placement recorded in the AI-extended ADR; incidents routed through the existing OT incident process, not a parallel AI one |
| Industrial security and functional safety standards (IEC 62443, IEC 61508 / 61511) | Zone separation, safety integrity levels, independence of protection layers, and management of change. | The architectural rule that AI is never a safety function, recorded as a pattern constraint in the pattern library |
| Wholesale market integrity rules (for example REMIT in the EU, market manipulation prohibitions elsewhere) | Prohibition on manipulative strategies, inside information handling, and records of trading decisions. | Strategy constraints in the agent authority boundary; decision records retained to the market standard |
| Consumer protection and supply obligations, including protections for vulnerable customers | Restrictions on disconnection, obligations to identify and support vulnerable households, and complaints handling. | Vulnerability handling recorded as a D1 justification and a control, with the detection path evidenced |
| Smart metering and data protection rules | Purpose limitation on consumption data, consent for granular readings in some regimes, and security requirements. | Data Lineage and Sensitivity Record for every meter data flow, with the granularity recorded |
| EU AI Act, where it applies | AI as a safety component in the management and operation of critical infrastructure appears in the high-risk list. | Regulatory Overlay Reference, mapped per operational system |
3. Calibrating the five dimensions
The dimensions do not change. What changes is what a 3 and a 4 look like when the subject matter is this sector, and which reading an assessor under delivery pressure reaches for first.
| Dimension | How to read it here | The mis-score to watch for |
|---|---|---|
| D1 Decision Consequence | Anything that can de-energize, energize, curtail, disconnect or delay restoration is 4. On the customer side, loss of supply to a household with a known vulnerability is 4. | Splitting the estate's anchors by department rather than by consequence, so the customer side scores itself against commercial impact. |
| D2 Autonomy | Control-room acceptance under alarm load is not review. Closed-loop remediation is 4 regardless of the operator's ability to intervene afterwards. | Scoring the designed intervention point rather than the one available during an event, which is when the system matters. |
| D3 Reversibility Deficit | Physical actions, market positions and disconnections are 4. Equipment damage and safety consequences do not reverse because the switch can be closed again. | Mechanism reversibility again: the breaker reclosing is not the same as the outage not having happened. |
| D4 Exposure and Scale | Network topology is the scale factor. A single bad instruction propagates along the network, and cascade potential should push D4 to 4 for anything on the transmission or primary distribution path. | Counting connected customers rather than network reach and interdependency. |
| D5 Sensitivity and Uncertainty | Operational telemetry is low sensitivity but often high uncertainty. Half-hourly consumption data is sensitivity 3. Take the higher sub-score. | Scoring OT systems as D5 = 1 because no personal data is involved, and losing the uncertainty half of the dimension entirely. |
4. One system, end to end
The overlay above is a map. This is one route across it: a single fault followed from detection to post-event review, with the record or control that attaches at each step.
5. What each gate adds
Additions only. Everything in the base gate definitions still applies; see the gate checklists for the common set.
| Gate | Sector addition | Why it is here |
|---|---|---|
| G1 | State which side of the OT boundary the system sits on and whether any output can become an operational instruction. Confirm no safety function is being delegated. | The answer determines the security regime, the review path and the tier floor, and it is expensive to change later. |
| G2 | Zone and conduit placement, protection independence, and the failure mode when the model is unavailable. Every operational data source needs a staleness limit, because stale telemetry is worse than none. | Availability and staleness are safety-relevant in a way they are not in most sectors. |
| G3 | Control-room procedure updated and rehearsed, kill-switch conditions defined for closed-loop systems, and an agreed maximum instruction rate. | Authorization has to include the human procedure. The system and the control room are one system. |
| G4 | Any change to autonomy, instruction scope or rate limits is Material. Seasonal reconfiguration counts as change. | Operational configuration drifts continuously, and it is the parameter that determines consequence. |
| G5 | Retention aligned to asset and incident investigation timescales, which are long, and to regulatory record-keeping for market activity. | Investigations into physical events reach back years. |
6. Controls and evidence worth adding
| Control | Where it attaches | Evidence it produces |
|---|---|---|
| Architectural prohibition: no AI in a safety function or protection path | Pattern library constraint; ADR conformance check | Pattern conformance recorded at G2, with any deviation escalated rather than self-certified |
| Instruction rate and scope limits for any acting system | Agent Card, enforced at the control layer | Machine-readable boundary comparable against actual instruction logs |
| Telemetry staleness monitoring with a defined maximum age | Data Lineage Record; runtime signal | Age-of-data metric per feed, with the action on breach stated |
| Operator acceptance-rate monitoring in the control room | Runtime signal; assurance cycle | Acceptance rate by shift and by alarm load, which is where it degrades |
| Vulnerability check before any disconnection or supply restriction | Control Matrix; workflow | Evidence the check ran and what it returned, per decision |
| Both-direction error monitoring for shutoff and curtailment decisions | Assurance cycle | Measured false-positive and false-negative rates, reported together |
7. Runtime signals to wire first
Control Plane onboarding order matters more than coverage in the first year (Chapter 18). These are the signals that earn their place earliest in this sector.
| Signal | Drift category | Suggested response |
|---|---|---|
| An AI service gains a network path into an OT zone it was not approved for | Security / integration | Immediate: this is a perimeter breach in the terms the OT regime already uses |
| Operator acceptance rate rises during high-alarm periods | Behavioral, with a classification consequence | Re-score D2 for the event condition, not the average condition |
| Telemetry feed exceeds its staleness limit | Data lineage and governance | Fail closed to the deterministic path; a model on stale data is worse than no model |
| Instruction volume or scope exceeds the approved envelope | Agent boundary divergence | Incident at any tier: suspend the acting path, preserve state, notify the control-room manager |
| Disconnection or debt-action rate changes by customer segment | Behavioral | Route to the customer vulnerability lead and the regulator-facing team together |
8. Failure modes this sector produces
Advisory in the control room, autonomous in the storm
The tell. During normal operation the operator considers each recommendation. During a major event, with three hundred alarms an hour, every recommendation is accepted.
The response. Score and instrument the event condition separately. If the system will be trusted absolutely during events, govern it as an acting system and put the protection independence in place accordingly.
The pilot that reached into the process network
The tell. A cloud analytics service is given a read path into the historian, and six months later a write-back capability is enabled to close the loop.
The response. Zone placement is an architecture decision recorded at G2. Enabling write-back is an autonomy change and a Major change, not a configuration flag.
Forecast confidence taken as certainty
The tell. Point forecasts are consumed by downstream optimization that has no representation of uncertainty, and positions or dispatch decisions are taken as if the forecast were fact.
The response. Require the uncertainty interval to cross the interface, and treat its absence as an assurance failure rather than a modelling nicety.
Collections automation meeting a vulnerable household
The tell. A debt model is efficient, accurate on its own terms, and disconnects a customer whose vulnerability flag was in another system.
The response. Make the vulnerability check a preventive control in the action path, not a report reviewed monthly.
9. A ninety-day start
If the sector is yours and the framework is new, this is the order that produces something defensible fastest. It assumes one part-time architect and one risk lead, not a programme.
- Split the estate into OT-side and customer-side and give each its own scoring session with its own anchors.
- List every system whose output can become an operational instruction, directly or through an integration. Include anything with a write path to a historian or SCADA layer.
- Confirm in writing, per system, that no safety function depends on a model. Where one does, that is the first remediation and it outranks all documentation work.
- Score the top three operational systems and the debt or disconnection system, with OT engineering and the customer vulnerability lead in the room.
- Set staleness limits on the operational feeds that matter and wire the alert.
- Write authority boundaries and kill-switch conditions for any closed-loop remediation.
- Measure control-room acceptance rate for a month, separating event periods from normal running.
- Run a G3 dry run on the highest-tier operational system with the control-room manager present.