On Call Journal
FeaturesLong read

What an Incident Response Drill Actually Tests (and What It Doesn't)

Drills test organizational coordination, not whether your monitoring actually works.

Columnist · · 9 min read · Updated
Cover illustration for “What an Incident Response Drill Actually Tests (and What It Doesn't)”
Features · October 2, 2026 · 9 min read · 2,077 words

Most engineering organizations have an incident response plan sitting in a wiki somewhere. Far fewer have tested whether that plan holds up when something is actually on fire. Mitigo Group's analysis of cyber fire drills makes this point directly: most organizations have incident response plans, far fewer know whether they work under pressure. The gap isn't negligence. Writing a plan and executing it while the pager is going off and the CEO wants an answer in ten minutes are two different skills, and a drill is built to train one of them, not both. When teams discover gaps mid-incident instead of during rehearsal, nobody knows who decides what, and that slows everything down at the exact moment speed matters most, so a contained technical problem turns into a business-wide one. None of this argues against running drills. It argues for being precise about what a drill actually closes, so teams spend money on the right rehearsal and the right tooling, not just more of the same rehearsal.

The human coordination layer that incident response drills genuinely stress-test

Drills earn their keep on the organizational side of incident response: who makes the call, who escalates to whom, who talks to customers and regulators, and how long each of those steps takes under pressure. A well-run drill exercises incident response decision-making under pressure, tests executive judgment when the facts on the ground are still incomplete, checks legal and communications readiness (notification obligations, pre-approved messaging templates, what gets said externally and when), and forces coordination across security, legal, HR, communications, and operations teams that otherwise rarely speak to each other during a crisis. It also sharpens time-to-containment and clears up who is supposed to escalate to whom. Most incident response failures start with unclear authority and slow decisions, and a drill is designed to rehearse exactly that layer.

The APCERT Cyber Drill 2026, titled "Incident Response: RMM-related Compromise," shows what this looks like at scale. Completed in August 2026 with 23 CSIRTs from 18 economies across the Asia-Pacific region taking part, the exercise had teams review and test their incident response procedures through simulated scenarios involving SQL injection, Kerberoasting, and other techniques, all inside a safe and controlled environment built around the misuse of Remote Monitoring and Management tools. That kind of cross-organizational rehearsal builds something real: organizations that have practiced a scenario make faster, steadier decisions during the actual event, because they're reacting to a structure they've already seen, even if the specific details differ. That value is genuine, but it is entirely organizational, which marks where the limit of a drill begins.

What a drill's controlled format structurally cannot test

Diagram: What a Drill Tests — and What It Doesn't. Visualizes: Visualize two parallel columns showing what incident response drills genuinely stress-test versus what they structurally cannot test.

A drill tests what participants agree to simulate. The observability stack, the alerting tools, and the detection mechanisms that would actually surface a real problem sit beneath that agreement, and a drill assumes they work instead of testing them. A tabletop exercise is discussion only: nobody touches a production system, and the team just talks through what it would do. A narrower operational drill might validate one link, like paging the on-call engineer or restoring a backup, but it still runs against a scenario someone built in advance. Even a full-scale simulation runs in a sandboxed or replica environment rather than against live production telemetry.

Drills are also not built to test firewall strength, malware detection accuracy, or whether a system can actually be exploited; that's what penetration testing is for. More relevant to engineering teams, drills don't test whether the telemetry and tooling the team would lean on during a real incident are producing signals that are accurate or even complete. The scenario arrives pre-labeled: the "incident" starts after someone has already decided what caused it. The hardest part of a real incident, noticing that something in production is behaving strangely in the first place, never happens inside a drill. A drill rehearses the response to an event that's already been named and defined, so it skips a much harder test: whether the team would have caught this on its own.

The detection problem: why production telemetry gaps are invisible inside a drill

Inside a drill, you never test the signals engineers depend on during a real incident, because the observability stack is never the patient. A drill skips past that question by answering it in advance.

Instrumentation gaps appear constantly in real incidents and almost never surface in drills. Teams find out, mid-outage, that the service that failed happens to be the one with no distributed traces, no structured logs, and a static threshold alert that fired once and then fell silent. Common instrumentation problems, managing the volume of telemetry data, keeping instrumentation consistent across a polyglot stack of services written in different languages, correlating traces across a complex architecture, and keeping the performance overhead of all that monitoring low, simply don't appear inside a synthetic scenario, because nothing in the drill is generating real telemetry under real load.

A drill built around a database outage can't reveal that the team's distributed traces would have been incomplete for that exact outage, because it never asks the observability stack to produce traces under real conditions. Post-deploy observability has the same blind spot. Most outages in modern systems come from a bad deployment, a configuration change, or a missing rollout guardrail, and drills typically skip this entirely because the incident is handed to participants already fully formed. CI passing is a necessary signal but not a sufficient one; production telemetry is the only real ground truth available during an incident, and a drill never checks whether that ground truth is readable at the moment someone actually needs it.

Alert noise as a precondition drills treat as solved

Every drill runs against one clearly labeled scenario. Real incidents occur inside an ongoing stream of alert noise that no drill replicates. The coordination skills a team practices in a drill are calibrated for a cleaner signal environment than the one on-call engineers actually work in.

Inside a drill, the alarm is by definition the one signal that isn't noise, because the facilitator has already told participants the incident has started. In production, the engineer who gets paged is already sitting inside a stream of deduplicated, grouped, still-imperfect alerts, and the mental load of sorting through that doesn't reset to zero just because a new page came in. That matters for how much a drill's lessons actually transfer: if the decision-making a team rehearses assumes a clean, unambiguous trigger, it may not hold up when the real first question is "is this an actual incident, or another noisy threshold firing again?"

Alert fatigue is a tooling problem before it's a people problem. A drill can train faster decision-making, but that doesn't fix the quality of the information those decisions get made on. The fix has to happen upstream: deduplication and grouping clean up the pipeline after the fact, but changing what gets instrumented at the source changes which signals get created, keeping the noise from entering the pipeline. A drill that includes an inject like "the monitoring stack is giving ambiguous signals" gets closer to what real incidents feel like, but most drills skip that condition entirely, leaving teams unprepared for one of the most ordinary things that happens during a real outage.

Tool fragmentation: the integration failure mode that a well-run drill conceals

A drill that goes smoothly inside a fragmented toolchain gives a team false confidence, because the places where tools connect, where trace context gets dropped between services, where alert routing is misconfigured, where two dashboards disagree about whether a service is actually up, are exactly where real incidents happen.

Tool sprawl looks like this in practice: too many disconnected monitoring tools generating duplicate alerts, multiple agents collecting the same telemetry in slightly different ways, thresholds that contradict each other depending on which tool set them. During a real incident, engineers end up jumping between dashboards, manually piecing together a timeline because each tool only sees its own slice of the system. The facilitator hands participants a coherent scenario, the team talks through what it would do, and nobody actually has to open three dashboards that disagree with each other about whether a service is healthy.

Teams use many tools, but the seams between those tools go untested until something breaks for real. Every extra tab an engineer opens during an outage is time the SLO keeps burning, so consolidating tools cuts down on exactly the kind of cognitive overhead that costs the most at the worst moment. The actual fix is production verification that works across the whole toolchain rather than inside each tool separately, so the integration points get checked continuously instead of being tested for the first time in the middle of an incident.

Agentic AI deployments as a new class of incident drills were not designed for

Standard incident response drills cover almost none of the failure modes that autonomous agents introduce: prompt injection, tools being misused by an agent acting outside its intended scope, cascading failures that spread across a multi-agent architecture. Teams shipping AI-generated code at the pace agents now allow are rehearsing for a different set of incidents than the ones they're likely to face.

Agentic applications involve routing agents, specialist agents, knowledge bases, external tool servers, and outside systems all interacting with each other, and that complexity makes debugging in production genuinely difficult without clear visibility into every step of the chain. Because agents frequently act without a human reviewing each individual step, a single failure can set off a chain of unauthorized or unintended actions before anyone notices. In a multi-agent setup, one compromised or malfunctioning agent can pass corrupted state downstream to the next agent in the workflow, and the problem compounds before a person is even in the loop. If a team's own incident response tooling runs on AI agents, those tools become part of the threat surface the team is supposed to be responding with, a recursive problem that standard drill design simply doesn't account for.

Competitive pressure is pushing teams to deploy agentic AI with minimal security review, often including unvetted components and code built through fast, iterative cycles. A growing volume of production code now ships whose behavior under failure conditions has never been checked. A drill that doesn't include an autonomous-agent failure scenario is already behind the threat model engineering teams are facing in 2026. The practical consequence is straightforward: AI-generated code shipped at agent speed needs automated, continuous verification in production as a standing safeguard built for a codebase that changes faster than human review can keep up with.

The tracking and coordination problem that keeps lessons from drills from compounding

A drill is only as valuable as the follow-through after it ends: whether the gaps it surfaces get tracked, assigned, and actually fixed. Most organizations don't have the infrastructure to guarantee that happens, and that includes the federal government.

A Congressional Research Service report on NLE 2026 found that the federal government has no formal system for tracking lessons learned across its exercises, and that multiple agencies run their own national preparedness drills without coordinating with each other, so lessons from one exercise may never reach another agency or feed into any shared strategy. NLE 2026, which began in April 2026, focuses on consequence management following critical infrastructure damage caused by cyberattacks tied to a transnational terrorist organization, and it's meant to test national preparedness and confirm the federal government can coordinate a response across agencies. The same CRS report notes that a prior Government Accountability Office review found FEMA had no formal mechanism to document and track best practices, lessons learned, and corrective actions from its own exercises.

Engineering teams replicate the same pattern at a smaller scale constantly. A postmortem turns up three instrumentation gaps and two places where escalation authority was unclear. Six months later, none of those items have an owner or a deadline attached to them, and the next incident rediscovers the same gaps from scratch. The structural fix is treating incident response improvement as a system with tracked inputs and outputs that gets revisited, rather than a recurring event that ends in a slide deck. Automated production verification closes a related loop from a different angle: instead of waiting for the next drill to rediscover that a telemetry gap still exists, it checks coverage against real production behavior on an ongoing basis, so the gap gets caught before the next incident finds it first.

Sources

  1. APCERT Cyber Drill 2026 “Incident Response: RMM-related Compromise”
  2. CRS Report Flags Gaps in Federal Disaster Exercise Oversight
  3. Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

More in Features