On Call Journal

Fix PR Generation as an Incident Response Pattern

Treating merged PRs as fixed until production confirms the regression vanishes.

Reporter · · 10 min read
Cover illustration for “Fix PR Generation as an Incident Response Pattern”
Autonomous Incident Diagnosis and Fix PRs · September 18, 2026 · 10 min read · 2,208 words

What "fix PR as the primary output" means in practice

Success looks like one specific thing here: a reviewed fix PR merges, and the regression it targeted disappears from production. Not the Slack thread going quiet. Not the postmortem doc getting filed away. The doc can wait. The regression can't.

Three stages make this work. Detection and root cause analysis correlate the telemetry and pin down the causal change, the actual commit or config shift that broke something. Fix generation turns that confirmed cause into a concrete code change and puts it up as a pull request. Then human approval and post-deploy verification close the loop: a person reviews and merges, and the system checks that the regression is actually gone once the fix hits production.

Merging a PR and confirming it worked are two different acts, and treating them as one is how teams end up "fixing" an incident that quietly keeps running underneath. The pattern isn't finished until production telemetry says so, not before.

A war-room thread full of updates is not the pattern. A generic alert summary is not the pattern. Neither is a runbook that restarts a service and calms the dashboards without touching the root cause, no matter how many indicators go green. An action item sitting in a postmortem that never ships isn't the pattern either. Those are all things that happen around a fix. None of them is a fix.

None of this removes the human from the loop, and it shouldn't. The approval gate at merge time carries the weight of the whole system. It's exactly because a person has to sign off that everything upstream of that moment needs to be fast and, more than that, accurate. A slow, sloppy RCA doesn't just delay the fix. It hands the reviewer a bad PR to approve under pressure, and that's worse than delay.

How good a fix PR can be depends on how good the RCA is

Following the chain backward, the whole pattern collapses into one dependency: a bad root cause analysis produces a wrong hypothesis, which produces a PR that either misses the actual bug or introduces a new one. The incident reopens, since it never really closed.

Anomaly detection alone doesn't perform the way vendor decks imply. A production benchmark testing single algorithms found precision landing between 52% (Z-Score) and 69% (trend analysis). An ensemble approach did better, around 70% precision, but still carried a 30% false positive rate. That's the honest floor, and any vendor claiming 95% noise reduction should get tested against it in a real proof of concept before anyone signs a contract.

Architecture is what actually moves the needle, more than the anomaly-detection layer people tend to fixate on. A multi-agent root-cause analysis system, deployed across 483 emergency incidents between October 2025 and March 2026, ran multiple agents chasing separate hypotheses in parallel instead of one agent working a single theory at a time. It correctly identified both the root cause service and the failure type in 82% of incidents, against 42% for a single-agent baseline on the same workload. That figure comes from an arXiv preprint, still provisional, not settled science. A gap that wide between 42% and 82% is evidence that how you structure the investigation matters more than which vendor's logo is on the dashboard. How you structure the investigation matters more than which vendor's logo is on the dashboard.

There are documented cases where an engineer acted on inaccurate guidance an AI agent had pulled from a stale internal wiki. The bug wasn't in the generated code. It sat in the context the agent trusted without question. Automated root-cause analysis is only as good as the documentation and telemetry it's allowed to read, and a stale wiki doesn't come with a warning label.

Two things follow from that. Telemetry freshness and documentation currency aren't nice-to-haves on a wishlist, they're prerequisites, full stop, and skipping them is how teams end up automating bad guesses faster. For the human at the merge gate, the PR description itself becomes the quality signal to watch. A description citing the specific log line, the specific commit, the specific metric that drove the diagnosis tells you the RCA did its job. One that just says "fixes the bug" tells you to slow down and look closer.

Diagram: Single-Agent vs. Multi-Agent RCA: The 40-Point Gap. Visualizes: Show a direct magnitude comparison between two root-cause analysis architectures tested across 483 emergency incidents (October 2025–March 2026): a single-agent baseline…

The tools that currently own different stages of the fix PR pipeline

The market has split into two camps: infrastructure investigation tools on one side, code-layer fix generation on the other. Nothing on the market runs the whole loop end to end yet, and treating either half as a complete solution is the mistake teams keep making.

One camp lives inside application code. Some tooling in this category, citing vendor-reported figures not independently checked, claims 38,000 or more issues fixed during a beta period at 94.5% root cause accuracy against internal benchmarks. Pricing as of January 2026 is a flat $40 per active contributor per month, unlimited usage, replacing an earlier per-fix credit scheme. The trust model here is explicit: teams can dial in how far the tool goes before a human has to step in, because not every org is ready to let an agent run unsupervised, and that's a reasonable place to land.

A second camp works the infrastructure side, letting admins define, through guardrails, exactly which remediation actions run automatically and which need sign-off first. An infrastructure investigation tool can tell you your system is broken, but it can't open a pull request into your application code the way a code-layer tool can. These aren't competing products so much as tools built for different halves of the same job, and buying one expecting it to do the other's work is a budget mistake as much as a technical one. The cost picture compounds fast once you're running both: per-host infrastructure fees in the teens to low twenties per month, per-host APM fees in the thirties to forties, tiered log billing, surcharges for custom metrics, and high-water-mark billing pegged to your 99th-percentile peak usage. Map this out before signing anything, not after the first invoice.

Resolve AI, founded by OpenTelemetry co-creators Spiros Xanthos and Mayank Agarwal, raised a substantial round at a valuation roughly eight times the funding amount from Lightspeed Venture Partners in February 2026, with total funding surpassing the raise by a further meaningful sum. Its multi-agent architecture chases hypotheses across infrastructure, code, and telemetry at once, tracing a failure across infrastructure, services, and code down to the exact commit that caused the regression. That gives it a wider reach than tools confined to application errors alone.

Chat-native platforms have entered this space too, correlating observability data with recent deploys to generate fix PRs scoped to the specific environment. One engineering team documented a case where the generated fix matched what a human engineer would have written, in 30 seconds instead of 30 minutes, a vendor-reported figure that reads as a best case rather than a typical one. Multi-agent platforms built around a context engine and a dedicated incident response agent can write fixes directly or hand them to a separate agent that drafts the pull request once a human signs off, on plans where that capability exists, again on self-reported numbers.

None of these tools closes the loop by confirming, in production, that the merged fix actually worked. That verification stage is either handled by hand or not handled at all, and this is the real gap in the market right now, not model quality. The pipeline behaves like a relay race missing its last leg: stages one through three, detection, RCA, and fix generation, are reasonably well served. Stage four, production verification, is where most implementations quietly stall out.

That's the gap One Patch is built to close, automatically correlating production telemetry against every merged PR, catching regressions that survived human review, and opening a follow-on fix PR when something breaks again, so confirming a fix held doesn't depend on someone remembering to check a dashboard three days later.

A real case shows why that fourth stage carries so much weight. Google Cloud declared gemini-2.5-flash-lite, the model behind an issue-summary feature, unavailable across several EU regions. The result: 80 to 90% of requests to the EU issue-summary endpoint started failing. The root cause wasn't the upstream outage itself. It was one missing guard clause in application code that didn't handle the upstream failure gracefully, no crash, no obvious red flag, just a high failure rate sitting quietly on one endpoint. The fix was small and bounded, a guard clause rather than a redesign, exactly the scope where an AI-generated PR earns trust. But knowing the guard clause got merged is a different fact from knowing the EU endpoint's failure rate actually came back down to baseline afterward. That confirmation step is precisely what's missing from most pipelines today.

Diagram: The Four-Stage Fix PR Pipeline — and Where It Stalls. Visualizes: Illustrate a four-stage relay race: Stage 1 Detection & Root Cause Analysis, Stage 2 Fix Generation (PR opened), Stage 3 Human Approval & Merge, Stage 4 Production…

Alert noise is what makes the fix PR pattern feel impossible before it feels inevitable

None of this works if the input is garbage, and for most teams right now, the input is garbage. The NeuBird AI State of Production Reliability and AI Adoption Report, drawing on more than 1,000 SRE, DevOps, and IT ops respondents, found 80% of organizations saying half or fewer of their alerts are actually actionable, and 77% of on-call teams fielding at least ten alerts a day. That's the daily reality most fixes are supposed to emerge from.

Toil rose as a direct consequence. Industry research has found a substantial share of SREs had considered leaving their role over on-call burden, with engineers fielding multiple pages per shift. Toil climbed to 30% of engineering time, up from 25%, the first increase in five years, despite 51% of organizations already running some form of AI tooling. Bolting automation onto a broken process just automates the noise faster, which should worry anyone betting on AI as a fix for burnout by itself.

The same report found 78% of organizations had experienced incidents where no alert fired at all, and 44% had outages traced back to alerts that were ignored or suppressed. Two failures wearing one face: too much noise, and silence exactly when it mattered most. Both trace back to the same broken instrumentation.

The fix PR pattern answers this structurally, not by asking engineers to somehow triage better at 2am. Route the alert to an agent first, before it reaches a human. The agent either produces a diagnosis with a proposed fix attached, or it returns a clean "no action needed" verdict. Either way, the engineer reviews a PR or dismisses a verdict; nobody stares at raw signal trying to guess whether it warrants action. Alert fatigue becomes something the tooling layer absorbs, not something a human is expected to power through on willpower and coffee.

Where this pipeline works, it doesn't just quiet the noise, it changes what the noise represents. Fixing the underlying issue before it can re-trigger is a different outcome than suppressing the alert that reported it. Teams running this well have described a meaningful shift away from operational toil and toward actual product work. That swing is the entire point of the exercise.

Governance that makes agent-generated fix PRs reviewable at speed

Generation used to be the bottleneck. Generation used to be the bottleneck, but it isn't anymore because complex, multi-file coding tasks now run significantly faster inside agent-assisted pipelines. Complex, multi-file coding tasks now run at something like 5 to 10x the speed inside agent-assisted pipelines. The constraint has moved downstream: verification is the bottleneck now, specifically getting a human to review and approve fast enough to keep pace with what the agents produce.

The baseline rule doesn't bend just because a PR came from an agent instead of a person: owner review, coverage thresholds, lint checks, static analysis, secret detection, all of it, no pilot-project exemptions because the code happened to be machine-written. Label every agent-generated PR with the tool and session ID behind it, something like agent:claude-code as a tag, so a security team can trace a PR back to the exact session that produced it once volume climbs into the hundreds or thousands of changes a month. That traceability stops being optional the moment you're past a handful of agent-authored PRs a week.

Frameworks like SOC 2 and ISO require independent review of every code change before production, and that requirement doesn't evaporate because an agent wrote the diff. But treating every agent PR as equally risky breaks velocity fast, and that's the wrong call. Tier by risk instead. A single-file guard clause or a routine dependency bump can move through a fast, automated-check lane. Anything touching multiple services, authentication logic, or a data migration needs full human review with a documented rationale attached, no shortcuts. Anthropic has described its internal approach in similar terms: a calibrated blend of automated agentic review and deterministic testing, with human attention reserved for the highest-risk changes rather than spread evenly across everything.

One problem in this space remains genuinely unresolved. Agents operating inside enterprise development environments often run with permissions no human engineer would survive an audit holding. Nobody has fully solved agent identity and scoped access yet. Any team adopting this pattern at scale should treat that as an open risk to manage actively.

Sources

  1. AI Incident Management Software: 2026 Evaluation Guide

More in Autonomous Incident Diagnosis and Fix PRs