Skip to content
CodeRabbit Agent for Slack

Your team's AI collaborator

One agent for your entire Software Development Lifecycle. Right in Slack.

Try for free

Get $50/user free agent minutes.

# eng-incidents
Datadog
Datadog 2:38 AM 🔴
Alert: checkout-service p99 latency > 10s
checkout-service p99 latency spiked from 380ms to 12.4s at 02:38 UTC.
Nadia K.
Nadia K. 2:41 AM
@CodeRabbit pull the Datadog traces for checkout-service and figure out where the bottleneck is.
CodeRabbit
CodeRabbit 2:41 AM
Pulled APM traces since 02:15 UTC. Every slow trace is stuck on the same downstream call, POST /api/inventory/reserve to inventory-service.

Checked GCP Cloud Run logs for inventory-service. Instance count dropped from 8 to 1 at 02:35 UTC. Autoscaler shows a config change pushed at 02:34 UTC that set maxInstances to 1. That came from PR #3301 merged earlier tonight. It was supposed to update env vars but accidentally included a Terraform change to the scaling config.
Marcus J.
Marcus J. 2:43 AM
Just saw the page. Nadia, you need a hand?
Nadia K.
Nadia K. 2:43 AM
CodeRabbit already found it. Accidental scaling change in #3301. @CodeRabbit open a revert for just the scaling config, don't touch the env var update.
CodeRabbit
CodeRabbit 2:44 AM
Done. PR #3315 opened: fix: revert inventory-service maxInstances to 12

Isolated just the scaling revert. The env var update from #3301 is untouched. Once merged, Cloud Run will scale back up within ~90 seconds.
Marcus J.
Marcus J. 2:45 AM
Merging. Nadia, watch the dashboard.
Nadia K.
Nadia K. 2:47 AM
Latency dropping. Back under 500ms. We're good.
Marcus J.
Marcus J. 2:48 AM
@CodeRabbit draft a postmortem and file it. SEV-2.
CodeRabbit
CodeRabbit 2:48 AM
Done. ENG-4120 created in Linear: "SEV-2 Postmortem: inventory-service scaling regression" with timeline, root cause, and remediation. Linked to PR #3301 and PR #3315.
BoldItalicUnderlineStrikethroughLinkOrdered listBulleted listBlockquoteCodeCode block
AddAaEmojiMentionVideoMicShortcuts