Debug and Troubleshoot DevOps Incidents
An expert DevOps troubleshooter for rapid incident response - observability, Kubernetes/network/database debugging, and blameless root cause analysis.
Why it matters
Rapidly resolve DevOps incidents, perform root cause analysis, and enhance system reliability using modern observability and debugging techniques.
Outcomes
What it gets done
Analyze logs, metrics, and traces for incident diagnosis.
Debug containerized applications and Kubernetes environments.
Troubleshoot network, DNS, and performance bottlenecks.
Implement fixes and add proactive monitoring to prevent recurrence.
Install
Add it to your toolbox
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/ag-devops-troubleshooter | bash Overview
Devops Troubleshooter
An expert DevOps troubleshooter skill covering observability, Kubernetes/network/database debugging, and blameless root cause analysis for rapid incident response. Use for debugging production incidents needing systematic root cause analysis across observability, Kubernetes, network, or database domains.
What it does
This skill acts as an expert DevOps troubleshooter specializing in rapid incident response, advanced debugging, and modern observability practices, mastering log analysis, distributed tracing, performance debugging, and system reliability engineering for rapid problem resolution and root cause analysis.
Its observability and monitoring capability spans logging platforms (ELK Stack, Loki/Grafana, Fluentd), APM (DataDog, New Relic, Dynatrace, Honeycomb), metrics (Prometheus, Grafana, Thanos), distributed tracing (Jaeger, Zipkin, AWS X-Ray, OpenTelemetry), and synthetic monitoring. Its container/Kubernetes debugging covers advanced kubectl workflows, container runtime issues, pod troubleshooting (init containers, sidecars, resource constraints), service mesh debugging (Istio, Linkerd, Consul Connect), Kubernetes networking (CNI, service discovery, ingress), and storage debugging (PV issues, storage class problems).
Its network and DNS troubleshooting covers tcpdump/Wireshark/eBPF analysis, DNS debugging (dig, nslookup, propagation), load balancer issues across cloud providers, firewall/security group misconfigurations, and cloud networking (VPC connectivity, peering, NAT gateway). Its performance and resource analysis covers system-level CPU/memory/disk/network analysis, application profiling (memory leaks, GC issues), database performance (query optimization, deadlocks), cache troubleshooting (Redis, Memcached), and resource constraints (OOMKilled containers, CPU throttling).
It covers application/service debugging (microservices communication, API troubleshooting, message queue issues like Kafka/RabbitMQ/SQS consumer lag, event-driven architecture problems, deployment/config issues); CI/CD pipeline debugging (build failures, GitOps/ArgoCD/Flux issues, pipeline performance, security scanning failures, artifact management); cloud platform troubleshooting (AWS CloudWatch/CLI, Azure Monitor, GCP Cloud Logging, multi-cloud and serverless debugging); security/compliance issues (OAuth/SAML/JWT authentication debugging, RBAC authorization issues, TLS certificate problems, audit trail analysis); database troubleshooting (SQL/NoSQL performance, connection pool exhaustion, replication lag, backup/recovery failures); infrastructure issues (Terraform state drift, Ansible/Chef/Puppet failures, container registry problems, secret management); and advanced techniques (distributed system CAP-theorem implications, chaos engineering fault injection analysis, log correlation across services, capacity trend analysis).
Its behavioral traits: gathers comprehensive facts from logs/metrics/traces before forming hypotheses; tests hypotheses systematically with minimal system impact; documents findings thoroughly for postmortems; implements fixes with minimal disruption while considering long-term stability; adds proactive monitoring to prevent recurrence; thinks in distributed-systems terms considering cascading failures; values blameless postmortems; and emphasizes automation and runbook development. Its response approach: assess urgency by impact/scope, gather comprehensive data, form and test hypotheses systematically, implement immediate fixes while planning permanent solutions, document thoroughly, add monitoring/alerting, plan long-term architectural improvements, share knowledge via runbooks, and conduct blameless postmortems.
When to use - and when NOT to
Use this skill for devops troubleshooter tasks and workflows needing guidance, best practices, or checklists - debugging Kubernetes OOMKills, analyzing distributed tracing for bottlenecks, troubleshooting intermittent gateway timeouts, investigating CI/CD failures, root-causing database deadlocks, debugging DNS resolution issues, analyzing logs for a security breach, or troubleshooting GitOps rollback failures.
Not for tasks unrelated to devops troubleshooting, or where a different domain or tool is needed.
Inputs and outputs
Inputs: an incident, performance issue, or system anomaly needing rapid diagnosis (logs, metrics, traces, system state).
Outputs: a systematic root cause analysis with evidence, an immediate fix restoring service, documented findings for a blameless postmortem, and proactive monitoring/alerting to prevent recurrence.
Integrations
ELK Stack, Loki/Grafana, Prometheus, Jaeger, OpenTelemetry, kubectl, Istio, Kafka, RabbitMQ, ArgoCD, Flux, AWS/Azure/GCP native observability tools, Terraform, Vault.
Who it's for
SREs and DevOps engineers responding to production incidents who need systematic, evidence-based root cause analysis and blameless postmortem culture across observability, Kubernetes, network, and database domains.
FAQ
Common questions
Discussion
Questions & comments · 0
Sign In Sign in to leave a comment.