Skill

Debug and Resolve System Errors

Systematic error analysis and resolution for production incidents - root cause identification, evidence-validated fixes, and preventive observability


75
Spark score
out of 100
Updated last month
Version 13.2.0

Add to Favorites

Why it matters

Systematically analyze and resolve errors across your application lifecycle, from local development to production incidents, to improve system reliability.

Outcomes

What it gets done

01

Investigate production incidents and recurring errors.

02

Perform root-cause analysis across services.

03

Design and implement observability and error handling improvements.

04

Propose fixes, tests, and preventive measures.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/ag-error-debugging-error-analysis | bash

Overview

Error Analysis and Resolution

An error analysis and resolution skill using systematic root-cause analysis and preventive observability improvements for production incidents. Use for investigating production incidents, root-causing recurring errors, or designing observability/error-handling improvements.

What it does

This skill acts as an expert error analysis specialist with deep expertise in debugging distributed systems, analyzing production incidents, and implementing comprehensive observability solutions - analyzing errors across the full application lifecycle from local development to production incidents using industry-standard observability tools, structured logging, distributed tracing, and advanced debugging techniques, with the goal of identifying root causes, implementing fixes, establishing preventive measures, and improving system reliability.

Its analysis scope adapts to whatever is provided - specific error messages, stack traces, log files, failing services, or general error patterns. Its instructions: gather error context, timestamps, and affected services; reproduce or narrow the issue with targeted experiments; identify the root cause and validate it with evidence; and propose fixes, tests, and preventive measures. Its safety constraints: avoid making production changes without approval and a rollback plan, and redact secrets and PII from shared diagnostics. It references a bundled resources/implementation-playbook.md for detailed analysis frameworks and checklists.

When to use - and when NOT to

Use this skill when investigating production incidents or recurring errors, performing root-cause analysis across services, or designing observability and error handling improvements.

Not for purely feature development tasks, when error reports/logs/traces cannot be accessed, or when the issue is unrelated to system reliability.

Inputs and outputs

Inputs: error messages, stack traces, log files, failing services, or general error patterns needing root-cause analysis.

Outputs: an evidence-validated root cause, proposed fixes and tests, and preventive measures/observability improvements - with production changes gated on approval and rollback plans.

Integrations

A bundled resources/implementation-playbook.md for detailed analysis frameworks and checklists; standard observability tooling (structured logging, distributed tracing).

Who it's for

Engineers investigating production incidents or recurring errors who need systematic, evidence-based root cause analysis and preventive observability improvements.

FAQ

Common questions

Discussion

Questions & comments · 0

Sign In Sign in to leave a comment.