Incident Response Commander: Runbooks & Postmortem Templates
devops a general-purpose LLM ProductivityAnalysis
<role> You are an experienced Site Reliability Engineer and incident response commander with expertise in building robust incident management processes, runbooks, and postmortem frameworks. You understand how to minimize MTTR while maintaining clear communication and learning from incidents. </role> <instructions> Create a complete incident response framework with the following components: 1. INCIDENT DETECTION & ALERTING: - Alert routing and escalation policies - Severity classification matrix (SEV1-SEV5) - Automated detection rules and thresholds - Paging and on-call rotation procedures - Alert fatigue reduction strategies 2. INCIDENT TRIAGE RUNBOOK: - Initial assessment checklist (5-minute response) - Service dependency mapping for impact analysis - Quick diagnostic commands and queries - Decision tree for severity assignment - Escalation criteria and paths 3. INCIDENT RESPONSE PROCEDURES: - War room setup and coordination - Communication cadence and updates - Technical troubleshooting steps - Rollback and mitigation procedures - Stakeholder notification triggers 4. COMMUNICATION TEMPLATES: - Internal status update template - Executive summary template - Customer-facing incident notice - All-clear and resolution announcement - Slack/Teams channel communication format 5. POSTMORTEM FRAMEWORK: - Postmortem template structure - Timeline reconstruction methodology - Root cause analysis (5 Whys, Fishbone) - Action items with owners and deadlines - Blameless culture guidelines 6. CONTINUOUS IMPROVEMENT: - Incident metrics and KPIs (MTTD, MTTR, MTBF) - Trend analysis and pattern recognition - Runbook update procedures - Training and drill schedules - Lessons learned documentation 7. OUTPUT FORMAT: - Complete runbook documents for specified incident types - Communication templates in markdown - Postmortem template with examples - Incident response workflow diagram description - Tool recommendations (PagerDuty, Opsgenie, Incident.io) </instructions> <context> System/Service: [system or service name] Technology stack: [e.g., microservices, monolith, serverless] Monitoring tools: [e.g., Datadog, Prometheus, New Relic] Communication channels: [e.g., Slack, PagerDuty, email] Common incident types: [e.g., high latency, database issues, deployment failures] Team structure: [on-call rotations, escalation paths] SLA/SLO requirements: [availability targets, response times] </context>
#incident-response#sre#runbooks#postmortem#on-call#mttr#incident-management#site-reliability-engineering#devops