Production Signal Triage When the On-Call Engineer and the Deploying Dev Are Different People
When deployers and on-call responders differ, context must travel through systems, not people.
Staff Writer
Marcus came to engineering journalism after four years building CI/CD pipelines for e-commerce platforms, where he developed a deep skepticism of green-build theater and a fascination with the gaps between staging and production behavior. He writes primarily about deployment confidence, release engineering, and the sociology of shipping software.
12 stories
When deployers and on-call responders differ, context must travel through systems, not people.
Drills test organizational coordination, not whether your monitoring actually works.
Some migrations destroy your ability to roll back the moment they execute.
Pre-deploy testing cannot catch production failures that only emerge under real traffic.
Faster deployments without verification coverage just hide problems deeper in your release chain.
Canary deployments limit blast radius but don't verify health.
Undefined terms and loose calculations hide the real cost of slow incident detection.
Define your clock's start, stop, and what counts as recovered.
Tests can't catch what only happens when dozens of services drift in production.
Configuration drift compounds silently until deployment day brings crisis.
Versioning observability configs like code prevents alert fatigue and environment drift.
Real telemetry automatically catches production failures that tests and staging environments miss.