On Call Journal

Real-Time Anomaly Detection on Production Metrics After Deploy

Catch regressions introduced by new deployments before they become production incidents.

Columnist · · 12 min read
Cover illustration for “Real-Time Anomaly Detection on Production Metrics After Deploy”
Telemetry Driven Regression Detection · September 11, 2026 · 12 min read · 2,646 words

Deploy a new service, and the monitoring question flips overnight. It's no longer "is something wrong right now?" It's "did this specific change cause something to go wrong?" That's a narrower question, and it needs a narrower method to answer it. General-purpose monitoring, the kind built on long-run historical baselines, absorbs the deploy event as just another data point in a rolling average. It never isolates the moment of change, so it can't tell you what the change did.

Deploy-correlated anomaly detection fixes this by treating the release as a hard boundary. Pre-deploy behavior becomes the comparison window. Post-deploy behavior becomes the thing under interrogation. Without that reset, a gradual regression introduced by a bad release just blends into background drift and never crosses any alarm. Latency climbing 5ms an hour, per OpenObserve's anomaly detection guide, never trips a static threshold on its own. Six hours later, it's an outage. That's the exact failure class deploy-scoped detection exists to catch, and understanding why it works means first understanding why the old method doesn't.

Why static thresholds fail at the deploy boundary specifically

Static thresholds fail in three distinct ways, and all three get worse at the exact moment a deploy lands.

First, context blindness. A 2% error rate during a Monday morning deploy window might be totally normal traffic churn. The same 2% at 4 AM is a page. A threshold can't tell the difference, because a threshold doesn't know what time it is or what just shipped.

Second, tuning fatigue. Whoever set the "right" number calibrated it against pre-deploy traffic patterns. The moment behavior shifts, which is exactly what a deploy is supposed to do, that threshold is stale. Nobody goes back and retunes it in real time, so it sits there, either too loose to catch anything or too tight to stop firing on noise.

Third, no adaptation to regime change. A deploy that doubles write volume, or introduces a new caching layer, or changes a downstream dependency, shifts what "normal" even means. The old threshold doesn't just become inaccurate, it becomes meaningless. It was built for an operating envelope that no longer exists.

Gradual degradations are the class of failure this breaks hardest on. A point-in-time threshold check has no memory of trajectory, so a slow climb toward failure looks identical to normal noise until the exact instant it crosses the line, and by then the damage is already compounding. Splunk's research (n=1,855) found 73% of organizations experienced outages linked to alerts that were ignored or suppressed, and the pattern behind a lot of those cases is thresholds misconfigured, alerts dismissed as noise, and the real signal buried in the pile. The deploy event is precisely the moment the operating envelope is most likely to shift, which makes it precisely the moment a static threshold is least trustworthy. So if thresholds don't work here, what does? The answer runs through baselines that reset at the deploy itself, built with models that can actually learn what normal looks like.

How deploy-correlated anomaly detection works mechanically

The core architectural decision isn't complicated to state, even though the implementation isn't trivial: use the deploy event as a baseline boundary, not just a timestamp on a graph. Hours or days of pre-deploy behavior become the training set for "normal." The window right after the deploy becomes the detection zone, and any deviation gets attributed to the change itself, not to seasonality or background drift. The baseline resets every time a deploy lands. The model isn't comparing today's traffic to last quarter's, it's comparing this hour to the hour just before the release went out.

A few algorithm families do the actual work here, per OpenObserve's guide. Statistical baselines, things like z-score and IQR, are fast and easy to interpret, and they're fine for stable metrics like cache hit rates. They fall apart on anything seasonal. Time-series forecasting models like Prophet or ARIMA forecast an expected value with a confidence interval, and flag post-deploy readings that fall outside the predicted band. Tree-based streaming models, Random Cut Forest being the notable one, handle continuous telemetry streams without needing a full batch retrain every time new data comes in, which matters a lot when deploys happen several times a day.

Establishing a usable baseline takes real data. OpenObserve's guide puts the minimum at 3 to 7 days of clean pre-deploy telemetry. Shorter than that, and the false positive rate climbs, because the model hasn't seen enough of what normal actually looks like across a full traffic cycle.

There's a fourth approach worth naming, from Microsoft Research and the University of Washington: a system called ARGOS, which uses LLM-generated anomaly rules as an intermediate representation. Multiple collaborative agents generate, validate, and deploy detection rules on their own, targeting three properties: explainability (an on-call engineer can actually read the rule and understand why it fired), reproducibility (same input, same output, every time), and autonomy (rules update as data distributions shift, without a human rewriting thresholds by hand). ARGOS outperformed prior state-of-the-art detection by up to 9.5% F₁ on public datasets, and by 28.3% F₁ on an internal Microsoft dataset. The paper doesn't explain why the internal gain was so much larger, but the number itself is notable. One of the motivating cases: a GPU training job running across 256 A100s hit a network hang that manually-written rules missed completely, because the signal was one that manually-written threshold rules missed entirely. That's a hard thing for a human to write a rule for, and an easier thing for a model trained on the shape of normal to catch.

None of this works without instrumentation feeding it. Deploy event markers have to reach the telemetry pipeline directly, or the system has no way to know when to reset the baseline window. Metric coverage has to span latency, error rate, throughput, and resource utilization, because partial coverage just creates blind spots that no amount of clever baselining can compensate for. And trace-level data has to be available for root cause attribution, because knowing something anomalous happened is a different problem from knowing which service caused it.

Progressive delivery as a structural amplifier for post-deploy detection

Route a slice of traffic to a new version, and the deploy stops being an event and starts being an experiment. That's the mechanical shift progressive delivery makes. Anomaly detection can now compare new-version metrics against old-version metrics running at the same time, on the same traffic mix, which is about as clean a signal as this kind of detection ever gets.

It also caps the damage. Catching a regression while it's affecting 5% of traffic is a fundamentally different event than catching it at 100%. The blast radius is bounded by the traffic weight assigned to the canary, full stop.

By 2026, the standard entry point for this is a canary starting around 5% traffic, run through tools like Argo Rollouts or Flagger, gated by automated metrics rather than a human eyeballing a dashboard (per the progressive delivery guide on dev.to). Promotion and rollback get driven by service level indicators directly. The deploy-scoped anomaly detection system sets the gate condition, and a human only gets pulled in when the results are genuinely ambiguous, not every single time a deploy ships.

The two practices need each other to actually work. Canary delivery without anomaly detection still shrinks the blast radius, but it still relies on someone noticing the degradation manually. Anomaly detection without canary delivery catches the regression, technically, but only after it's already hit every user. Put them together, and the operating question changes shape entirely: not "did we survive the deploy?" but "what did the telemetry show at 5% exposure?" That's a question a system can answer on its own, without a war-room Slack thread and six people staring at the same graph.

Where alert noise enters the picture and how deploy-scoping reduces it

Alert noise post-deploy tends to break into three categories. False positives come from alerts firing against a stale or poorly calibrated baseline, one that doesn't reflect what normal actually looks like after the change shipped. Redundant notifications come from the same underlying regression tripping five different alerts, each one technically accurate and each one adding to the pile. Over-sensitive triggers come from thresholds set before the deploy changed the operating envelope, firing on deviations too minor to warrant action.

Deploy-scoping addresses the first category structurally, not incidentally. Because the baseline window is drawn from pre-deploy behavior specifically, post-deploy seasonality and traffic shifts don't contaminate what the model considers "normal."

Alert correlation and compression, the AIOps-style approach, can cut raw event volume by as much as 90%, according to incident management KPI research. But compression alone doesn't solve attribution. Fewer alerts that still can't be traced to a specific deploy just means fewer things to ignore, not more clarity.

The cost of getting this wrong isn't abstract. Toil rose 30% in 2025, the first increase in five years, and 70% of SREs surveyed by Catchpoint in 2025 indicated that on-call stress impacted burnout and attrition. In the same 2025 incident management report, financial services teams extended outages by hours because real alerts got dismissed as noise, a direct consequence of alert systems that can't separate deploy-caused anomalies from background variance. The fix isn't more alerts or fewer alerts in isolation. It's a baseline that resets at deploy time, producing alerts that are fewer in number and sharper in attribution, so an engineer who sees one knows immediately it maps to a specific PR, not to some vague drift nobody can explain.

What makes a post-deploy anomaly actionable rather than just observable

Detecting an anomaly is necessary, but it's not the finish line. An anomaly score sitting on a dashboard, waiting for a human to diagnose it, route it, and figure out what to do about it, still produces toil. It's just toil with better math behind it.

Turning detection into action takes a full pipeline. Detection flags the deviation inside the post-deploy window. Attribution links that deviation to the specific deploy event, and where trace data supports it, to the service or component most likely responsible. Triage groups related signals so the team gets one incident record instead of fifteen separate alerts pointing at the same root cause. Resolution then follows one of two paths: an automated runbook executes, or a fix PR gets opened for an engineer to review, rather than a Slack thread asking if anyone knows what's going on.

The speed difference here is substantial. A root cause analysis draft that would take an engineer 2 to 3 hours to put together manually can be ready 5 minutes after an incident starts, with AI-assisted platforms, per OpenObserve's guide on AI incident management. Organizations running AIOps report a 62.1% reduction in mean time to resolution and, more tellingly, an 81.7% decrease in repeat incidents, according to a 2025 IJERET study. The repeat-incident number matters more than the resolution-time number, because it reflects the system actually identifying the correct root cause, not just responding faster to the same recurring problem.

Human review still belongs in the loop for anything high-risk. Low-risk playbooks, service restarts, straightforward rollbacks, can run automatically. Infrastructure changes still need an engineer's sign-off. The right end state of a post-deploy anomaly detection system is a reviewed fix PR sitting in someone's queue, not a war-room escalation pulling six people off their work. That's the difference between reactive firefighting and structured review, and it changes what being on-call actually feels like.

OnePatch sits directly in this part of the pipeline. It verifies every pull request against live telemetry after it deploys, catches regressions before anyone has to get paged, and opens fix PRs on its own when something breaks, built specifically around this deploy-scoped, actionable-response model rather than a generic alerting layer bolted on afterward.

What it takes to instrument a team's stack for deploy-scoped detection

Four things have to be in place, and skipping any one of them leaves a gap that no amount of clever modeling can paper over.

Deploy event markers need to reach the telemetry pipeline directly. Without a reliable signal at the moment of promotion, the system has no way to reset the baseline window or attribute an anomaly to a specific change. That means CD pipeline integration isn't a nice-to-have, it's the thing that makes the whole approach possible. Inferring a deploy from a shift in traffic shape is a guess. A structured signal at the moment of promotion is not.

Metric coverage has to match the actual surface area of what can fail. Three anomaly types matter here: metric anomalies (latency, error rate, throughput), log anomalies (new error patterns the change introduced), and trace anomalies (spans that got slow and weren't slow before). A gap in instrumentation is a gap in detection, plainly. If a deploy degrades a downstream dependency that isn't instrumented, that regression simply doesn't show up, no matter how good the anomaly model is.

Baseline data needs real depth, 3 to 7 days minimum of clean pre-deploy telemetry, per OpenObserve's guide. Teams shipping multiple times a day run into a real tension here: short windows between deploys mean the system has to handle overlapping or rapidly succeeding baseline periods, and that requires logic built for it, not an afterthought bolted on.

Alert routing has to connect to the deploy context itself. An anomaly should land with the engineer who owns the PR that caused it, not with a generic on-call queue that then has to figure out whose problem it is. Assigning ownership to the team responsible for the change cuts time-to-acknowledge and kills the "is this even real?" debate that eats up the first twenty minutes of a lot of incidents.

Agentic development adds a wrinkle here that's still new enough to be unsettled. AI-generated code shipping at agent speed means deploys can outpace the baseline windows meant to stabilize around them, which forces a choice: shorten the baseline window and accept more false positives, or build continuous retraining that adapts to rapid succession deploys. The 2025 AI Agent Index, a joint effort from MIT, Cambridge, Stanford, and Harvard Law, found that developers disclose far less about safety practices than they do about capabilities. That same gap shows up here in operational terms: agentic systems shipping code into production without automated verification behind them are running with real observability blind spots, whether or not anyone's tracking that as a governance issue yet.

How to evaluate whether a detection system is actually working post-deploy

Three properties, borrowed from the framing Microsoft Research used for ARGOS, tell you whether a production anomaly detection system is doing its job or just generating noise with better branding.

Explainability comes first: can an on-call engineer read why an anomaly fired, not just that it did? If the rule is a black box, nobody can assess it, and nobody can improve it when it's wrong. Reproducibility comes second: the same input metrics need to produce the same result every time. A non-deterministic alarm burns engineer trust fast, and trust, once burned, is expensive to rebuild. Autonomy comes third: the system has to adapt when data distributions shift, because a new deploy that changes the operating envelope shouldn't require someone to manually reconfigure thresholds every time.

Beyond those three properties, watch two trends over time. Alert volume should decline while incident detection rate holds steady or improves, that's compression without losing coverage, not compression that's just hiding real signal. And time-to-attribution, the gap between an anomaly firing and knowing which deploy caused it, should trend down consistently.

If alert fatigue keeps rising despite all this instrumentation, despite the baseline resets, despite the correlation and the routing, that's the tell. It means the system is generating signal, technically, but not signal anyone can act on, which is the exact failure mode deploy-scoped detection was built to close.

Sources

  1. AI Anomaly Detection: Complete Guide for DevOps & SRE 2026
  2. arxiv.org
  3. openobserve.ai
  4. automationanywhere.com

More in Telemetry Driven Regression Detection