[AI Data Center Security] “AI Data Center Physical Security: 4 HAZOP-Style Questions a Checklist Can’t Answer”

AI Data Center Physical Security: 4 HAZOP-Style Questions a Checklist Can’t Answer
By Jin Ho Kang · OT/ICS Security Consultant
Table of Contents
- Why AI Data Center Physical Security Has Outgrown the Checklist
- What HAZOP Actually Does That a Checklist Doesn’t
- Cyber-PHA: The Bridge Between Process Safety and AI Data Center Physical Security
- Applying HAZOP Guide Words to Cooling, Power, and Physical Access
- What This Actually Costs When Nobody Asks the Questions
- Bringing It Together
- References
Most physical security audits follow a well-worn pattern: a binder, a checklist, and a walk around the facility confirming that each box can be ticked. Cameras present – check. Badge readers functioning – check. Fence line intact – check. What a checklist almost never asks is the question a process safety engineer asks by habit: what happens when this specific system deviates from its intended behavior, and what does that deviation cascade into? That question is the entire basis of HAZOP, a technique built for chemical plants, and it turns out to be exactly the question AI data center physical security needs answered as GPU halls start combining extreme power density, liquid cooling, and hyperscale physical footprints in ways no static checklist was designed to anticipate. This piece is about what changes when AI data center physical security borrows that question instead of borrowing another checklist.
Why AI Data Center Physical Security Has Outgrown the Checklist
The numbers behind data center reliability make the case on their own. Uptime Institute’s 2025 Annual Outage Analysis found that power-related issues remain the single most frequent hazard to data center uptime, and cooling failure is the next most common cause behind them – and the report explicitly ties the pressure on both systems to soaring AI demand, which is straining infrastructure designs built for a different power and thermal profile (Uptime Institute, 2025). More than half of surveyed operators, 54%, said their most recent significant outage cost over $100,000, and one in five put that figure above $1 million (Uptime Institute, 2025). Perhaps most telling for anyone thinking about AI data center physical security specifically: the share of outages attributed to human error from failure to follow procedure rose by ten percentage points in a single year (Uptime Institute, 2025). A checklist confirms that a procedure exists. It does not tell you whether that procedure will be followed correctly under pressure, or what happens to the facility if it isn’t – which is precisely the gap AI data center physical security teams are now being asked to close, and the gap a HAZOP-style review is built to find before an incident does.
What HAZOP Actually Does That a Checklist Doesn’t
HAZOP – Hazard and Operability Study – is a structured, team-based technique, standardized under IEC 61882, that identifies hazards by systematically applying guide words to each parameter of a system rather than checking whether a control is present (Wikipedia; Tractian). The standard guide words are simple on their face: No/None, More, Less, As Well As, Part Of, Reverse, and Other Than (Tractian, 2025). A multidisciplinary team walks through each system node and asks what it would mean for flow to be reversed, for temperature to be more than intended, for a signal to arrive as something other than expected – then works through the causes, consequences, and existing safeguards for each deviation (Primatech). Originally developed at Imperial Chemical Industries in the 1970s for chemical plants, HAZOP has since been adapted to nuclear, oil and gas, and increasingly to safety- critical control systems of every kind, precisely because the guide-word method does not depend on the specific hazards of any one industry (Wikipedia, “Hazard and operability study”). Applied deliberately, this is what turns a generic physical security review into AI data center physical security engineering: instead of asking “is there a cooling system,” the question becomes “what happens if flow through this cooling loop is less than intended, and what safeguard catches it before it becomes a thermal event.”
Fig. 1 – The 7 HAZOP guide words applied to AI data center physical security nodes
Cyber-PHA: The Bridge Between Process Safety and AI Data Center Physical Security
HAZOP’s reach into security specifically runs through a related technique called Cyber PHA, or cyber HAZOP – a cybersecurity risk assessment methodology for industrial control systems that follows the same deviation-based logic as a traditional process hazard analysis, formalized in ISA-TR84.00.09 and aligned with ISA/IEC 62443-3-2 (Wikipedia, “Cyber PHA”; exida). The method exists because a conventional PHA already documents exactly the information a security assessment needs – worst-case consequences, existing protection layers, and severity rankings for every credible deviation – so a Cyber PHA reuses that foundation and asks an additional question for each scenario: could the initiating event, or one of the control barriers meant to stop it, be triggered by a cyberattack rather than a mechanical fault (exida, 2025; ISA InTech, 2020)? For AI data center physical security, that reframing matters enormously, because the systems that keep a GPU hall cooled and powered – building management systems, chiller controllers, UPS management interfaces – are cyber-physical systems in exactly the sense Cyber PHA was designed for, and a compromise of the control system can produce the same physical consequence as a mechanical failure, just triggered on an attacker’s schedule instead of a random one. This is, in effect, AI data center physical security and industrial cybersecurity converging into a single discipline rather than two adjacent ones.
Applying HAZOP Guide Words to Cooling, Power, and Physical Access
Run the guide words against three systems that define AI data center physical security today and the value becomes concrete rather than abstract. On the liquid cooling loop that increasingly serves high-density GPU racks, “Less” flow at a distribution manifold is the deviation that precedes thermal throttling or an outright thermal event, and the relevant safeguard question is whether flow-rate deviation triggers an automatic load shed before temperature limits are crossed – not whether a cooling system is merely present. On the electrical side, “Reverse” applied to a backup generator’s transfer switch surfaces a classic backfeed hazard during maintenance, a scenario well understood in industrial electrical safety but easy to miss in a facility built around uninterruptible power rather than rotating generation. On physical access control, “Other Than” applied to a badge event is the guide word that generates tailgating and social-engineering scenarios a badge-system checklist simply cannot surface, because the checklist only confirms the reader works – it never asks what happens when a credential is used in a way other than its intended single-person, single-direction pattern. This is where CPTED principles and building management system design experience compound the value of the method: someone who has actually configured a BMS/BAS integration or laid out a CPTED-informed site plan can populate the “causes” and “existing safeguards” columns of a HAZOP worksheet with specifics a generalist security auditor has no way to know, which is exactly the kind of AI data center physical security review a checklist was never built to produce. Without that domain depth, an AI data center physical security assessment tends to default back to the checklist it was supposed to replace.
Fig. 2 – Applying HAZOP guide words to three AI data center physical security systems
What This Actually Costs When Nobody Asks the Questions
The financial case for treating AI data center physical security as a deviation-analysis problem rather than a checklist problem is already in the outage data. With power issues the leading cause of outages at 52% and cooling failure close behind, and with 54% of significant outages costing over $100,000 and one in five costing over $1 million, the population of incidents a HAZOP-style review is designed to catch is not a hypothetical tail risk – it is the majority of what is already going wrong (Uptime Institute, 2025). The ten-percentage-point rise in procedure-related human error outages is arguably the more important number for AI data center physical security specifically, because it shows that having a documented procedure – the thing a checklist verifies – is not the same as the procedure surviving contact with an actual deviation under time pressure. That gap between “a control exists” and “the control holds when the parameter it governs actually deviates” is precisely the gap HAZOP and Cyber PHA were built to close, and it is a gap that scales with the stakes involved, which for a hyperscale AI training facility are considerably higher than for a typical enterprise data hall.
Fig. 3 – The outage data behind the case for HAZOP-style AI data center physical security reviews
Bringing It Together
A checklist tells you whether the parts of a facility are present. HAZOP and its security-focused sibling, Cyber PHA, tell you what happens when any one of those parts deviates from its intended behavior – mechanically, procedurally, or because someone made it deviate on purpose. For AI data center physical security, where cooling and power systems are being pushed harder than the infrastructure they were designed around and where a single-day outage can cost seven figures, that distinction is not academic. It is the difference between an audit that confirms a system exists and an engineering review that tells you exactly how it fails – and that engineering-grade review is what AI data center physical security, at hyperscale, actually needs.
References
- Wikipedia, “Hazard and operability study”
- Tractian, “HAZOP: Definition, Guide Words, and Steps” (2025)
- Primatech, “HAZOP (Hazard and Operability) Study”
- Wikipedia, “Cyber PHA”
- exida, “ISA/IEC 62443 Cybersecurity Services”
- ISA InTech, “Cyber-Related Process Hazard Analysis” (2020)
- Uptime Institute, “Annual Outage Analysis 2025” (press release)
- Data Center POST, “Maintaining Data Center Uptime: Critical Challenges and Solutions in 2025”