On Call Journal

Incident Management Roles and Responsibilities in On-Call Engineering

Clear roles before the alert fires prevent chaos when production breaks at 2 AM.

Features Editor · · 9 min read
Cover illustration for “Incident Management Roles and Responsibilities in On-Call Engineering”
On Call Toil and Alert Noise Reduction · September 22, 2026 · 9 min read · 2,044 words

Incidents cost businesses an estimated $700 billion, and the damage rarely comes from the outage itself. It comes from the confusion around it: three engineers debugging the same log file, or an hour passing with nobody telling customers a thing. When an alert fires at 2 AM, the first question every responder asks is "is this mine?" That question needs an answer before the pager goes off, not during the scramble that follows. Clear role clarity resolves an incident in twenty minutes rather than letting it drag into a postmortem titled "what went wrong, besides the obvious.""

Atlassian's incident response documentation names the failure mode directly: incidents get worse when responders can't communicate, can't cooperate, and don't know what everyone else is working on. That failure is visible in two opposite shapes, caused by teams' incident response habits, and most teams only ever notice the one that burned them last. Over-coordination happens when everyone converges on the same problem at once, three people staring at the same dashboard while nobody watches the door for the CEO's message on a messaging app. Under-coordination happens when everyone assumes someone else has it, and the alert sits acknowledged but untouched for twenty extra minutes. Both waste the same resource, and it's the one resource nobody gets back.

The incident commander: what it means to own the whole incident without touching the code

The incident commander owns the response, not the incident, and that distinction trips up more organizations than it should. Usually a member of the IT or DevOps team, the IC defines the incident action plan and drives decision-making once things start going sideways. Asana frames the job as seeing the big picture of the incident and breaking it into manageable pieces small enough for individual engineers to act on. Skip that decomposition, and defects, system errors, and miscommunication become easy to compound rather than resolve.

What the IC refuses to do matters as much as what they take on. They don't debug, and they don't write the fix. The job is holding the map while everyone else navigates the terrain, and inexperienced ICs get this backward constantly, reaching for the keyboard when they should be reaching for the phone. An IC who starts fixing things has stopped coordinating them, and nobody else in the room picks up that job while they're distracted.

The work breaks into four stretches. Before an incident happens, the IC builds communication channels, drafts a general action plan, and trains the team, work that reads as overhead until it's missing and everyone's improvising. At the moment an incident starts, the IC reads the situation and adapts the standing plan into a step-by-step response specific to what's actually broken. Once the incident is underway, the job turns into delegation: identifying which engineers are needed, pulling them in, and prioritizing by severity and business impact rather than by whoever's shouting loudest in the channel. As the incident evolves, the role shifts again into something closer to facilitator, unblocking stuck responders and stepping into communications when that channel breaks down.

Primary and secondary on-call engineers: the responders who touch production

The primary on-call engineer is first contact with a live problem. The job is initial response: acknowledge the alert, confirm the problem exists, and start mitigation before the situation compounds. A primary who spends fifteen minutes building the perfect diagnosis while the service stays down has made the wrong trade, every time, no exceptions. Speed resolves the incident. Elegance can wait for the postmortem.

Secondary on-call exists for incidents that outgrow the primary's expertise. Secondaries handle the complex cases, the ones needing architectural judgment or carrying real business stakes, and they typically end up leading the post-incident review once the fire's out. The handoff from primary to secondary is where role clarity either earns its keep or fails. It saves time only when the escalation trigger, the person authorized to make that call, and the context handed to the secondary at pickup are all clear. Miss any one of those three. The confusion just moves one level up the chain instead of disappearing.

Being on call isn't free, and treating it as background responsibility misreads what it costs. Engineers typically put somewhere around 30 to 40 percent of their working bandwidth toward on-call responsibilities during a rotation, a third to two-fifths of an engineer's week redirected from building toward standing guard. That time takes priority over whatever sat on the roadmap that week, regardless of what the roadmap owner would prefer.

The communications lead: keeping stakeholders informed without pulling responders off the problem

The communications lead manages every message flowing in and out of the incident, internal and external both, so the people fixing the problem never have to stop and answer "any updates?" from six different directions. That means writing status updates for very different audiences (an engineering peer needs different information than a senior business stakeholder does) and fielding questions that inevitably come from leadership, customers, and sometimes the press.

Atlassian's model puts status page ownership on the communications manager too, with collecting and triaging customer responses as a secondary duty. Status page ownership is one of the most commonly dropped balls in incident response: someone assumes it's automated, someone else assumes it's covered, and it sits stale for forty minutes while customers refresh a page looking for an answer that isn't coming.

The role goes by other names: communications officer, communications lead, and some organizations split off a separate social media lead to handle public channels. Smaller teams often skip the dedicated role entirely and let the IC absorb it, and that decision carries a real cost rather than a convenience. Every stakeholder question the IC answers personally is time not spent coordinating the technical response, the one job nobody else in the incident can do instead.

Supporting roles that complete the structure: tech lead, scribe, and problem manager

Diagram: The Hidden Tax on Engineering Time. Visualizes: Visualize the breakdown of how engineering time is consumed, contrasting building versus operational burden.

The tech lead, sometimes called the on-call engineer or subject matter expert, is the senior technical responder building theories about what's broken and deciding what changes get made. The IC holds the coordination; the tech lead holds the diagnosis. The two work in close partnership because neither can do the other's job well under pressure. The tech lead also documents key theories and actions as the incident unfolds, raw material for the postmortem later, and pages in additional subject matter experts when the problem crosses outside their own knowledge. Atlassian's model lets an incident manager appoint multiple tech leads at once when more than one workstream is active, a practical fix for incidents spanning three services that need three separate diagnostic threads running in parallel.

The subject matter expert is a narrower role: a responder who knows the specific system failing, often because they built it or currently own it. SMEs suggest and implement fixes, give the team context, and pull in other experts as needed. The distinction from tech lead comes down to scope. The tech lead coordinates the technical response across the whole incident, while SMEs get called in for depth on one particular piece of it.

The scribe records the incident as it happens: the timeline, who did what, when decisions got made. It's the role most likely to get skipped, and its absence causes the most damage weeks later, once the postmortem starts. Without a scribe, the postmortem becomes a reconstruction from memory, and memory under incident pressure runs unreliable, and often quietly self-serving, since people tend to remember their own actions more favorably than a timestamped log would show.

The problem manager, also called the root cause analyst, picks up where the fix leaves off. The job is to run and record the postmortem, identify what actually caused the incident, and track the remediation tickets that come out of it. This role bridges one incident and the systemic fix that keeps it from happening again, and skipping it leaves a team fighting the same fire repeatedly, surprised every single time.

Where roles break down under pressure: the friction points

The handoff between IC and tech lead is the most fragile boundary in the whole structure, and it deserves more attention than it usually gets. The IC has to delegate real technical authority to the tech lead without losing track of what's happening, and when that boundary blurs, one of two things follows. Either the IC starts micromanaging the technical response and slows it down, or the IC checks out and loses the coordination thread that justified having an IC.

The primary-to-secondary escalation fails in a different way but just as often, and the common fix makes it worse. Teams that define the escalation trigger purely by time elapsed miss the incidents where a secondary's architectural knowledge would have changed the outcome had it arrived sooner. The trigger needs to track the signal. Patience is not a diagnostic tool.

Communications is where the lack of a single owner becomes hardest to miss. Research on incident response tooling has found that 83 percent of organizations juggle four or more separate tools during a live incident, and in that kind of fragmented environment, every other role ends up bleeding into communications by default: answering a message on a chat tool here, updating a status page there. Coordination collapses under the weight of that context-switching because nobody was assigned to catch it.

The scribe gap follows a similar pattern in smaller teams, where scribing tends to stay informal or simply doesn't happen. What follows isn't a postmortem so much as a negotiation over what people think happened.

The effect of alert noise on role-defined teams

Clean role definitions assume a manageable signal, and that assumption falls apart fast in practice. Roughly 80 percent of organizations report that half or fewer of their alerts are actually actionable, and 77 percent of on-call teams field at least ten alerts a day. Role clarity was never built to solve a signal problem of that size, and pretending otherwise just hides where the real cost sits.

For the primary on-call engineer, that volume changes the job in a way no role description accounts for. If less than half of what's firing deserves a response, the primary spends real cognitive effort triaging noise before incident response even starts. That's a tax the org chart never shows; it appears instead in burnout numbers that are easy to recognize and hard to quantify precisely, caused by the primary's need to triage noise before incident response even starts.

The predictable response to that volume is suppression: muting alert categories, raising thresholds, silencing whole channels. It's a rational move against the noise, and it creates its own risk in the same motion. Some research puts the figure at 44 percent of organizations tying outages directly to alerts that were ignored or suppressed; other research puts it at 73 percent. Whichever number a given team trusts, the underlying cause is the same: a defense mechanism against volume causes it to occasionally swallow the one alert that actually mattered, and it does so often enough that this is a pattern rather than an edge case.

The time cost of on-call toil and what it takes from engineering teams

Added together, these numbers describe an industry spending an outsized share of its best hours on maintenance instead of progress. Engineers reportedly burn around 40 percent of their time on incident management rather than building the product they were hired to build, the aggregate cost of exactly the role friction and alert noise described above. According to State of Incident Management research, 78 percent of developers spend 30 percent or more of their time on manual toil, and operational toil itself climbed to 30 percent from 25 percent, the first increase in five years.

Close to a third of engineering capacity gets redirected from what an organization actually wants to build toward keeping the lights on, and no framework erases that cost. An IC who coordinates without touching code, a tech lead who diagnoses without losing the thread to leadership, a communications lead who owns the status page so nobody else has to: none of that makes the toil disappear. What it does is split toil that resolves from toil that compounds, and that split is the entire argument for taking these roles seriously.

Sources

  1. Role of an Incident Commander: Real-Time Crisis Control [2026] • Asana
  2. Understanding incident response roles and responsibilities | Atlassian
  3. On-Call Scheduling & Alerting | Incident Management Platform | Atlassian
  4. runframe.io

More in On Call Toil and Alert Noise Reduction