READINESS & PLANNING
How to Build an OT Incident Response Plan
A practical guide to building an incident response plan that actually works for operational technology — roles, escalation, asset understanding, evidence sources, and how to keep it current.
Most OT incident response plans that don't work in practice fail for the same reason: they were written by adapting an IT incident response template, filling in OT-sounding language, and never testing whether the plan holds up against a realistic scenario. A plan that hasn't been exercised against something like "a compromised engineering workstation is discovered on a Friday afternoon" is a document, not a capability.
This is a practical guide to building a plan that's actually usable — grounded in what an OT incident response engagement typically needs to move fast, not in what looks complete in a binder.
Start With What You're Protecting
Before writing procedures, define what matters most if it's compromised or unavailable: which processes are safety-critical, which are production-critical, and which systems those processes actually depend on. This sounds obvious, but many plans skip straight to generic procedures without ever establishing this baseline — which means the plan can't help anyone prioritize during an actual incident.
Roles and Escalation
Define, by name or role, who has authority to make containment decisions, who represents engineering and operations, who handles communications, and who escalates to leadership. Critically, define escalation paths for outside business hours — a plan that only works Monday to Friday during office hours isn't a real plan, since incidents don't respect that schedule.
- Incident commander — has authority across both IT and OT decisions, or a fast path to someone who does
- Engineering/operations lead — understands what each affected system currently does in the process
- IT/security lead — handles technical containment, investigation, and evidence preservation
- Communications owner — manages internal and external updates on a defined cadence
- Executive escalation contact — informed early, not after major decisions are already made
Understand Your Assets Before You Need To
An incident is the worst time to discover you don't have an accurate asset inventory. At minimum, know what engineering workstations, HMIs, historians, jump servers, and SCADA components exist, what they run, and who's responsible for them. This doesn't need to be a perfect, continuously-updated CMDB to be useful — even a reasonably current spreadsheet beats discovering assets for the first time during a live incident.
Network Architecture and Segmentation
Document what's actually connected to what — not what the original design intended, but what's true today, including the exceptions and workarounds that accumulate over time. Segmentation between IT and OT, and between OT zones, directly determines how far an incident can spread and how effectively it can be contained. If segmentation has degraded since it was designed, the plan needs to reflect that reality, not the original architecture diagram.
Know Your Evidence Sources
Identify in advance where useful evidence would actually come from if you needed it: which systems log meaningfully, where those logs are retained and for how long, which engineering workstations hold project files and configuration backups, and what network monitoring — if any — exists at the IT/OT boundary. Knowing this ahead of time is the difference between preserving evidence quickly and discovering, mid-incident, that critical logs rotate out after 24 hours.
Communications Planning
Plan communications templates and cadences in advance for at least three audiences: internal operational teams who need to know what's safe to do, executive leadership who need accurate status without technical noise, and any external parties — regulators, insurers, customers — who may need to be informed depending on severity. Decide this before an incident, not while drafting the first update under pressure.
Vendor and OEM Contacts
Maintain a current list of vendor and OEM contacts for critical control system components, including emergency contact paths outside normal business hours. When an incident touches proprietary equipment, the ability to reach the right vendor contact quickly — rather than starting with a generic support line — can materially change how fast recovery moves.
Isolation Decision Points
Rather than a blanket "isolate affected systems" instruction, define decision points: for each major system category, what does isolation require operationally, who has to approve it, and what's the safe sequence to do it in? This is where engineering input during plan development matters most — these decisions are much harder to make correctly for the first time during a live incident.
Backup and Recovery Planning
Document what's actually backed up, how often, where backups are stored, and — critically — whether they've ever been tested by actually restoring from them. An untested backup is a hypothesis, not a recovery plan. Include engineering workstation project files and PLC logic backups specifically; these are frequently missed in backup strategies built around server infrastructure.
Exercise the Plan
A plan that's never been tested against a realistic scenario will have gaps you can't predict from a desk review. Tabletop exercises — walking through a hypothetical scenario like a compromised engineering workstation or a ransomware event reaching OT-adjacent systems — surface unclear roles, missing contacts, and unrealistic assumptions while the stakes are still hypothetical.
Capture Lessons Learned
After any real incident or exercise, document specifically what worked, what didn't, and what changes as a result. This step gets skipped constantly, usually because everyone is relieved the incident is over and wants to move on. Skipping it means the same gaps get rediscovered the next time, under worse conditions.
Keeping the Plan Alive
Plans go stale as fast as the environment changes — new systems, new vendors, personnel turnover in key roles. Review and update the plan on a defined schedule, not "whenever someone remembers." A plan that reflects last year's architecture and contact list is a false sense of security.
Building a plan this way takes real effort, and it's reasonable to want outside input — particularly for the asset understanding, isolation decision points, and exercise design, where experience with real OT incidents matters. Our Incident Readiness Assessments and OT Tabletop Exercises are built specifically to test and strengthen a plan like this before it's needed for real.
ABOUT THE AUTHOR
OTR³ — OT Incident Response & Industrial Cyber Recovery, focused on critical infrastructure across the GCC. Learn more about OTR³ →
RELATED SERVICES
Incident Readiness Assessments
Identify risk, validate controls and strengthen your operational resilience.
Learn MoreOT Tabletop Exercises
Realistic OT/IT scenarios to test plans, people and decision-making.
Learn MoreIncident Response Retainers
On-demand access to senior OT cyber experts when you need them most.
Learn MoreRELATED INDUSTRIES
RELATED ARTICLES
Why Industrial Organizations Need an OT Incident Response Retainer
Why establishing OT incident response access before an incident occurs saves critical time — and what pre-incident familiarization, escalation paths, and readiness integration actually provide.
Learn MoreVendor Remote Access in OT: Incident Response When Third-Party Access Is Compromised
How to scope, contain, and safely re-enable vendor and third-party remote access into OT environments after a suspected compromise — VPNs, jump hosts, shared credentials, and session review.
Learn MoreReady to talk to OTR³?
Active incident or planning ahead — reach out and we'll point you to the right engagement.

