CodeRabbit Agent for Slack
Your team's AI collaborator
One agent for your entire Software Development Lifecycle. Right in Slack.
Try for freeGet $50/user free agent minutes.
# eng-incidents
Datadog 2:38 AM 🔴
Alert: checkout-service p99 latency > 10s
checkout-service p99 latency spiked from 380ms to 12.4s at 02:38 UTC.Nadia K. 2:41 AM
@CodeRabbit pull the Datadog traces for checkout-service and figure out where the bottleneck is.
CodeRabbit 2:41 AM
Pulled APM traces since 02:15 UTC. Every slow trace is stuck on the same downstream call,
Checked GCP Cloud Run logs for
POST /api/inventory/reserve to inventory-service.Checked GCP Cloud Run logs for
inventory-service. Instance count dropped from 8 to 1 at 02:35 UTC. Autoscaler shows a config change pushed at 02:34 UTC that set maxInstances to 1. That came from PR #3301 merged earlier tonight. It was supposed to update env vars but accidentally included a Terraform change to the scaling config.Marcus J. 2:43 AM
Just saw the page. Nadia, you need a hand?
Nadia K. 2:43 AM
CodeRabbit already found it. Accidental scaling change in #3301. @CodeRabbit open a revert for just the scaling config, don't touch the env var update.
CodeRabbit 2:44 AM
Done. PR #3315 opened:
Isolated just the scaling revert. The env var update from #3301 is untouched. Once merged, Cloud Run will scale back up within ~90 seconds.
fix: revert inventory-service maxInstances to 12Isolated just the scaling revert. The env var update from #3301 is untouched. Once merged, Cloud Run will scale back up within ~90 seconds.
Marcus J. 2:45 AM
Merging. Nadia, watch the dashboard.
Nadia K. 2:47 AM
Latency dropping. Back under 500ms. We're good.
Marcus J. 2:48 AM
@CodeRabbit draft a postmortem and file it. SEV-2.
CodeRabbit 2:48 AM
Done. ENG-4120 created in Linear: "SEV-2 Postmortem: inventory-service scaling regression" with timeline, root cause, and remediation. Linked to PR #3301 and PR #3315.