Free incident report template
Document an incident clearly: what happened, impact, timeline, root cause, actions taken and steps to prevent it. Sample report included.
What's in this template
INCIDENT REPORT · INC-2026-0318 Checkout outage Report by Arjun Menon, Site Reliability Lead · 6 October 2026 |
Date of incident | Saturday 3 October 2026 | Duration | 47 minutes (11:52 to 12:39 IST) |
Severity | SEV-1 | Status | Resolved |
Systems affected | Checkout service, mobile app | Incident lead | Arjun Menon |
Summary
Customers of Kirana Basket could not complete orders on the website or app for 47 minutes during Saturday's peak shopping hour. A database setting changed in a routine update capped the number of connections, so checkout requests timed out.
Impact
6,240 Failed checkout attempts | ₹18.6 lakh Estimated lost orders | 1,130 Support tickets raised |
No customer data was lost or exposed. No payments were taken for failed orders.
Timeline (IST)
Time | Event |
|---|---|
11:40 | Database configuration update deployed |
11:52 | Checkout error rate rises above 40%; on-call alerted |
12:05 | Incident declared SEV-1; war room opened |
12:21 | Connection limit identified as the cause |
12:31 | Configuration rolled back |
12:39 | Checkout success rate back to normal; incident resolved |
Root cause
The update reduced the database's maximum connections from 500 to 50, copied by mistake from a test environment file. Our pre-release checks did not compare configuration values between environments.
Actions taken
- Rolled back the configuration at 12:31
- Posted status updates on the app banner and X every 15 minutes
- Sent an apology with a ₹100 voucher to all affected customers on 4 October
Prevention
- Add an automatic check that blocks config changes that differ from production by more than 20% · Sana Qureshi · 20 October
- Alert when database connections exceed 80% of the limit · Arjun Menon · 16 October
- Freeze non-urgent changes on weekends from 10:00 to 22:00 · Vikram Bhatt · 13 October
- Share this report with all engineering teams · Arjun Menon · done
Lessons learned
Our monitoring caught the problem quickly, but we spent 16 minutes looking at the wrong service. A clearer runbook for checkout errors would have halved the outage.
Find it in: Incident report template