Chapter 07 / 07

Resolution

Fix deployed, SLOs recovering, post-mortem

Step 1Fix deployed

The config change that resolved the incident is documented here. Tracefox links the deployment event back to the open incident automatically — no manual ticket chasing required.

Executive viewUnderstand what a fast, well-documented recovery looks like — and how Tracefox reduces the time between detection and resolution.

Recovery story

2.6h
Time to recover
10:48:22 UTC
Alert cleared
Final burn rate

A single config change — db connection pool size increased — stopped the error cascade. Tracefox linked the deployment event to the open incident automatically, closing the loop without a manual update. The SLO returned to target within 2.6 hours of the fix landing.

Fix deployed

DB connection pool size increased

Increased the HikariCP connection pool size from 10 to 25 in payment-service. The pool had been exhausted under peak checkout load, causing all incoming requests to queue and eventually time out.

Alert cleared

Alert resolved at 10:48:22 UTC
SLO burn rate dropped below the alert threshold — on-call acknowledged and closed.

SLO recovering

After the fix was deployed, the SLO burn rate fell from above-alert levels down to 1× baseline over 2.6 hours. The SLO window is no longer in breach and the error budget has begun recovering. No further action is required.

Post-mortem

Post-mortemresolved
TimelineIncident started ~2024-03-15 09:21 UTC. Alert cleared 2024-03-15 10:48 UTC. Total duration: ~2.6h.
Root CauseTimeout waiting for connection from pool after 30000ms (pool.size=10, pool.active=10)
FixIncreased the HikariCP connection pool size from 10 to 25 in payment-service. The pool had been exhausted under peak checkout load, causing all incoming requests to queue and eventually time out.
PreventionRunbook updated. Alert threshold reviewed. Config change tracked in deployment record.
OwnerOn-call engineer assigned. Post-mortem review scheduled within 48 hours.