On Call Journal

The On-Call Engineer's Checklist for Verifying AI-Generated Code After It Ships

AI-generated code fails in ways your monitoring wasn't built to catch.

Contributing Editor · · 10 min read
Cover illustration for “The On-Call Engineer's Checklist for Verifying AI-Generated Code After It Ships”
Automated Production Verification After Every PR · October 5, 2026 · 10 min read · 2,217 words

CloudBees's 2026 survey found that 92% of enterprise technology leaders said they trust AI-generated code is ready for production. The same survey found that 81% of those leaders said production issues tied to AI-generated code had risen. Those two numbers describe the same population of respondents, and together they say something simple: the gates these teams trust (code review, static analysis, staging environments, a green CI run) are passing code that still breaks once it meets real traffic. Reviewers are not careless, and AI models are not unreliable in some vague sense. A model infers what code should do from statistical patterns in its training data and the prompt it's given. None of that context lives in a pull request, so none of it factors into what the model writes.

That gap produces failures that pre-deploy checks are structurally unable to catch. Code review won't catch either of these, because both look correct in isolation. Static analysis won't catch them, because both are syntactically and logically sound. Staging won't catch them either, because it rarely reproduces the concurrency or the data volume that triggers the failure.

There's a second problem layered on top of the context gap: the same system that writes the code is often the system that writes the tests validating it. A passing test suite in this setup confirms internal consistency, not correctness. A passing test suite confirms an AI-generated change is consistent with itself, because there was never an independent party involved to verify it against anything else. It says nothing about whether it will survive contact with the system it's about to join.

How AI code fails differently from human code

AI-generated code doesn't fail randomly. It fails in a handful of recognizable ways, because current monitoring stacks were built around the failure patterns of human-written code, not this.

The first pattern is the hallucinated API call: code that calls a method that doesn't exist, or that exists but behaves differently than the model assumed. It compiles, and linting passes, because the method name is plausible and the syntax is valid. The failure appears at runtime, when the call either throws or, worse, silently returns something you didn't expect. Industry research identifies this as one of the leading defect classes in AI-generated code, and it's a failure type that simply doesn't occur in human-written code at the same rate, because a human engineer who isn't sure an API exists will usually go check.

The second pattern is logic drift: an implementation that looks reasonable, passes a basic test, and quietly solves a slightly different problem than the one specified. This doesn't trip an alert, because nothing about it looks wrong until the wrong number shows up in someone's account.

The third pattern is tests that prove nothing, written by the same agent that wrote the code they're meant to check. The test suite turns green without the underlying problem going away.

The fourth pattern is coverage that skews toward the happy path. Models optimize for the scenario the prompt describes, which tends to be the common case, not the race condition, the partial rollback, the concurrent access pattern, or the timeout. Those are underrepresented in training data relative to how often they occur in production, so they're underrepresented in the code a model writes to handle them.

The deeper issue connecting all four patterns is that existing dashboards and alert thresholds were built around failure modes someone anticipated. AI-generated code introduces failure surfaces nobody on the team has conceived of yet, so no threshold is waiting, no runbook is written, and no dashboard panel points at the right metric. An on-call engineer is often looking at data that was built to show them some other, anticipated failure, not this one.

Diagram: Four Failure Patterns Unique to AI-Generated Code. Visualizes: Visualize the four named, sequential failure patterns that AI-generated code produces, as identified in the article: (1) Hallucinated API call — compiles and lints, fails at…

The on-call engineer's verification job starts after the PR merges

Every verification control most engineering organizations currently rely on, code review, linting, SAST, staging, lives before the deploy. Once the merge happens and the code reaches production, the on-call engineer becomes the only person still positioned to catch what slipped through, and without a deliberate structure for doing that, they're set up to lose that fight by default.

CloudBees's 2026 survey quantifies where that accountability lands when things go wrong: 46% of respondents said the CTO or VP of engineering is ultimately accountable for AI-related production failures, and 32% said accountability falls to the engineering lead or team associated with the tool that produced the faulty code. In practice, accountability becomes operational at the moment an alert fires at 2am, and the person holding the pager is the one who has to translate an abstract accountability structure into an actual diagnosis, in real time, often with less context than anyone who touched the code during the day.

Part of what makes that job harder than it should be is a feedback loop that doesn't close. An agent writes a patch, the PR merges, and the agent's session ends there. The next time that agent (or a different instance of the same model) touches related code, it starts from zero runtime knowledge of how its last change actually performed. That's a structural reason fixes to AI-introduced bugs often take multiple redeploy cycles instead of one: nothing in the loop remembers what happened last time.

The fix for this isn't another pre-deploy gate. The verification surface needs to shift to where the ground truth actually lives, which is live production telemetry, not another round of static checks against code that hasn't run yet.

This is the specific gap between a merged PR and what that PR actually does once it's live, and it's the problem One Patch is built to sit inside. It automatically compares telemetry gathered after a deploy against the baseline recorded before that deploy went out, so the on-call engineer has a structured signal already waiting, built from what the system actually did.

Scope and instrumentation checks to run before anything else

The first question to answer after an AI-generated PR ships is whether what actually deployed matches what was reviewed, and whether a failure in that code would even be visible if one occurred. Two checks come before everything else.

Gate 1 is scope verification. Confirm that the deployed diff matches the ticket, the prompt that generated it, and the file set that was actually approved. A diff that drifts outside its stated boundary is a signal worth stopping on before anything downstream of it gets trusted.

Gate 2 is an error instrumentation audit. Confirm that errors in the new code path are actually being recorded somewhere a person will see them. AxonBuild's study of AI-generated applications found that 17 of 21 third-party apps recorded errors nowhere a human would ever look: a user hit a failure, and it vanished without a trace. Run through the specific places this tends to happen: catch blocks that swallow an exception instead of logging it, error logs missing a request ID that would let anyone trace the failure back to a specific user or session, error paths that emit a generic message with none of the upstream state that caused it, and log statements present on the happy path but absent from the failure branch sitting right next to it. This audit has to happen before anything resembling a production smoke test, because there's no value in testing for failures a system can't record.

A third check belongs in this same early phase: deploy gate confirmation. Verify that something actually checked this push before it reached production. AxonBuild's analysis found that at least 17 of the 21 third-party AI-built apps it audited had nothing checking a push before it went live, and some of those apps had switched off their own type checks and lint rules at build time. Confirm that type checking, linting, and static analysis weren't quietly disabled somewhere in the build pipeline, and make sure you have a documented rollback path ready before you need it under pressure.

Production smoke tests and error-rate segmentation by code origin

Once scope and instrumentation have been confirmed, the next layer of the checklist moves into active verification against live traffic.

Gate 4 is production smoke testing: running tests against real traffic patterns and live dependencies, not against mocks or seed data. If a smoke test on one of these paths shows a policy violation or a regression, that's a trigger to validate the specific change immediately, before the failure has a chance to cascade into something that touches more of the system.

Gate 5 is error-rate segmentation by code origin: tracking errors per request separately for AI-authored changes and human-authored changes, rather than folding both into a single aggregate error rate. Segmented, it becomes visible and attributable. This also does something useful beyond the immediate incident: when error patterns tied to AI-generated code are tracked and attributed rather than lost in aggregate noise, that information can inform both the on-call engineer handling the current incident and the next agent session that touches related code, partially closing the runtime feedback gap described earlier.

This is the exact comparison One Patch automates. After each deploy, it checks live telemetry against the baseline recorded before that deploy and flags regressions tied to the specific code change responsible for them, so the on-call engineer is looking at a diff of what changed in production behavior rather than interpreting a raw dashboard from scratch.

Instrumentation gaps specific to agentic and multi-agent workloads

Agentic and multi-agent systems introduce instrumentation gaps that standard distributed tracing doesn't cover automatically, and these gaps matter more as more production code runs through agent pipelines.

The first gap sits at the Model Context Protocol boundary. When an agent calls out to an MCP server, the agent produces one trace and the server produces a separate trace, and in many implementations there's no context propagation connecting the two by default. Without this wired up, an on-call engineer looking at a trace sees a tool call that took an unusually long time and returned some blob of data, with nothing showing what the MCP server actually did inside that window.

The second gap is sub-agent visibility. If the sub-agent is where a bug actually originated, nothing in the standard trace points there.

The third gap is async context propagation. HTTP propagates trace context automatically as part of the request. None of these three gaps are hypothetical edge cases. They're the default behavior of the tools most teams are already running agentic workloads on top of, and they need to be checked explicitly.

Keeping alert noise from AI-generated code from drowning real signals

A rise in AI-generated production failures means a rise in the number of alerts those failures generate, and tuning existing alerts isn't enough to keep that volume from burying the signals that actually matter. The fix has to happen further upstream, at the point where a signal either does or doesn't get created.

All three categories can be eliminated without losing any real signal, and deleting these alerts outright, rather than tuning their thresholds, is the single change with the largest effect on mean time to acknowledge.

A three-layer framework describes where alerts should land. The page layer is reserved for critical SLO burns that need immediate human action, so it should make up only a small fraction of total alert volume. The log layer holds record-only signals with no action attached, and this is where most alert volume belongs and where most of the noise currently lives when it hasn't been sorted properly.

You should treat every new page-level alert that an AI-authored change generates as provisional, not trusted on sight. New code introduces failure modes nobody has validated an alert against yet, so a new alert tied to new code shouldn't be granted the same standing in the rotation as an alert that's been tested against real incidents over months. One Patch's approach to this problem works by catching regressions immediately after deploy and opening fix PRs on its own, which shrinks the number of incidents that ever reach the pager, addressing alert volume at the point where it's generated.

Agentic safety controls that must be in place before AI-generated code reaches production

OWASP's Secure Coding with AI guidance treats several categories of agentic behavior as distinct attack surfaces worth controlling before any of this code reaches production: MCP servers acting as tool providers, rules files that persist instructions across sessions, CI/CD systems an agent can modify, and the permissions and sandboxing governing what an agent is allowed to touch. Each of these maps directly onto a failure mode described earlier in this piece. If an agent has unrestricted access to CI configuration, it can disable the type checks and lint rules that Gate 3 depends on. An MCP server with no scoped permissions can be called in ways nobody reviewing the original PR anticipated.

The controls that close these gaps are specific rather than general: restrict what an agent's CI/CD access can modify, separate from what a human's access can modify. None of these controls replace the post-deploy checklist described above. They narrow how much damage a misconfigured or compromised agent can do before that checklist ever gets a chance to run. That is why they belong in place before the first line of AI-generated code reaches a production system.

Sources

  1. OnePatch - Automate on-call

More in Automated Production Verification After Every PR