Skill

Respond to Incidents with SRE Expertise

SRE-grade incident response: severity triage, observability-driven investigation, communication cadence, and blameless post-mortems.


79
Spark score
out of 100
Updated last month
Version 13.1.0

Add to Favorites

Why it matters

Act as an expert incident responder with deep Site Reliability Engineering (SRE) knowledge. This asset provides guidance, best practices, and actionable steps for rapid problem resolution, effective communication, and comprehensive post-incident analysis.

Outcomes

What it gets done

01

Assess incident severity and impact on users and business.

02

Establish incident command and communication channels.

03

Investigate incidents using observability and SRE techniques.

04

Implement resolution strategies and validate recovery.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/ag-incident-responder | bash

Overview

Incident Responder

An SRE-grade incident response framework covering first-five-minutes triage and incident command, observability-driven investigation, a fixed communication cadence, minimal-viable-fix resolution, and blameless post-mortems with P0-P3 severity classification. Use when running or improving incident response - triage, investigation, communication, or post-mortem process; skip it for tasks outside incident response.

What it does

Acts as an incident response specialist with deep Site Reliability Engineering (SRE) expertise, combining urgency with precision across modern incident management - from the first five minutes of an incident through resolution, recovery, and blameless post-mortem.

When to use - and when NOT to

Use this skill when working on incident responder tasks or workflows that need guidance, best practices, or checklists. Do not use it for tasks unrelated to incident response, or when a different domain or tool is actually needed.

Inputs and outputs

Immediate actions in the first 5 minutes: assess severity and impact (user count, geographic spread, revenue/SLA impact, blast radius); establish incident command (a single Incident Commander, a Communication Lead, a Technical Lead, and a war room with communication channels); and stabilize immediately via traffic throttling, feature flags, circuit breakers, rollback assessment of recent deployments, and resource scaling.

Investigation draws on observability tooling (distributed tracing via OpenTelemetry/Jaeger/Zipkin, metrics via Prometheus/Grafana/DataDog, log aggregation via ELK/Splunk/Loki, APM, real user monitoring) and SRE techniques (error-budget burn-rate analysis, change correlation against the deployment timeline, dependency mapping, cascading-failure analysis of circuit breakers/retry storms/thundering herds, and capacity analysis).

Communication follows a fixed cadence: internal status updates every 15 minutes during an active incident, technical detail for engineering, business-impact summaries for executives, and external status-page updates, support briefings, and regulatory notification where required. Documentation standards require a timestamped incident timeline, decision rationale, impact metrics, and a full communication log.

Resolution follows a minimal-viable-fix-first approach with risk assessment, staged rollout, and validation, followed by recovery validation across SLIs, real user monitoring, performance metrics, and dependency health. Post-incident process covers the first 24 hours (continued monitoring, resolution communication, data collection, team debrief) and a blameless post-mortem (timeline analysis, root-cause analysis via five whys/fishbone diagrams, contributing factors, action items, and follow-up tracking), feeding into system improvements across monitoring, automation, architecture resilience, and process.

Severity is classified P0-P3: P0/Critical (complete outage or security breach, <15min acknowledgment, <1hr resolution, updates every 15 minutes with executive notification) down to P3/Low (cosmetic issues, next-business-day response, <72hr resolution via standard ticketing).

SRE best practices covered: error budget management and burn-rate-driven feature freezes, reliability patterns (circuit breakers, bulkhead isolation, graceful degradation, exponential-backoff retries), and continuous improvement via MTTR/MTTD tracking and a blameless learning culture. Tooling spans PagerDuty, Opsgenie, ServiceNow, and Slack/Teams for incident management, plus unified dashboards and automated runbook diagnostics for observability.

Response principles: speed matters but accuracy matters more, communicate frequently at the right technical depth, fix first and understand root cause later, document everything, and treat every incident as a chance to improve.

Who it's for

SRE teams and incident commanders running structured incident response - from first-five-minutes triage through communication cadence, observability-driven investigation, and blameless post-mortem - who want a repeatable framework rather than ad hoc firefighting.

FAQ

Common questions

Discussion

Questions & comments · 0

Sign In Sign in to leave a comment.