Skill

Engineer Production Observability Systems

Designs production observability systems: monitoring, distributed tracing, log management, SLIs/SLOs, and alerting.

Works with prometheusgrafanainfluxdbdatadognew relic

80
Spark score
out of 100
Updated 11 days ago
Version 15.7.0

Add to Favorites

Why it matters

Establish and maintain robust, production-grade observability systems for enterprise applications, ensuring high reliability and performance through comprehensive monitoring, logging, and tracing.

Outcomes

What it gets done

01

Design and implement monitoring, logging, and tracing infrastructure.

02

Define and track Service Level Objectives (SLOs) and alerting strategies.

03

Investigate and resolve production reliability and performance issues.

04

Optimize observability pipelines for cost and efficiency.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/ag-observability-engineer | bash

Overview

Observability Engineer

Designs production-grade observability systems covering monitoring, distributed tracing, log management, SLI/SLO frameworks, and alerting. Use when designing monitoring/tracing/logging systems, defining SLIs/SLOs, or investigating production reliability issues.

What it does

Provides expert guidance for designing production-grade monitoring, logging, tracing, and reliability systems for enterprise-scale applications - covering SLI/SLO definition, alerting strategy, distributed tracing, and cost-optimized observability architecture.

When to use - and when NOT to

Use this skill when designing monitoring, logging, or tracing systems, defining SLIs/SLOs and alerting strategies, or investigating production reliability or performance regressions. Not a fit for a single ad-hoc dashboard, situations without access to metrics/logs/tracing data, or application feature development rather than observability work. Follows a four-step process: identify critical services, user journeys, and reliability targets; define signals, instrumentation, and data retention; build dashboards and alerts aligned to SLOs; validate signal quality and reduce alert noise - while avoiding logging sensitive data or secrets and balancing alert coverage against noise.

Inputs and outputs

Covers monitoring and metrics infrastructure: Prometheus/PromQL and recording rules, Grafana dashboard design, InfluxDB retention policies, DataDog/New Relic/CloudWatch enterprise monitoring, and high-cardinality metrics handling. Distributed tracing/APM guidance covers Jaeger, Zipkin, AWS X-Ray, OpenTracing/OpenTelemetry instrumentation, service mesh telemetry (Istio/Envoy), and correlating traces/logs/metrics for root cause analysis.

Log management guidance covers the ELK Stack, Fluentd/Fluent Bit, Splunk, and Loki, plus structured logging, retention policy design, and security/compliance log analysis. Alerting and incident response guidance covers PagerDuty routing/escalation, Slack/Teams notification workflows, alert correlation and noise reduction, runbook automation, on-call rotation management, and blameless postmortems.

SLI/SLO management covers defining and measuring SLIs, establishing SLOs, error budget/burn rate calculation, SLA compliance monitoring, and chaos engineering integration for proactive reliability testing. OpenTelemetry guidance covers collector deployment, auto-instrumentation, sampling strategy, and vendor-agnostic multi-backend export. Infrastructure monitoring covers Kubernetes (Prometheus Operator), container metrics, multi-cloud monitoring, database performance, and network/CDN/storage monitoring.

Chaos engineering guidance covers Chaos Monkey/Gremlin fault injection, circuit breaker monitoring, disaster recovery testing, and RTO/RPO validation. Dashboard guidance covers executive vs. operational dashboards, custom Grafana plugins, multi-tenant access control, and automated reporting. Observability-as-code guidance covers Terraform/Ansible for monitoring infrastructure and GitOps-based dashboard/alert management. Cost optimization covers retention policy tuning, sampling rate adjustment, and tool ROI analysis. Enterprise guidance covers SOC2/PCI-DSS/HIPAA compliance monitoring, SAML integration, and ITSM tool integration (ServiceNow, Jira Service Management). AI/ML integration covers anomaly detection, predictive capacity forecasting, automated root cause correlation, and intelligent alert clustering.

Integrations

Spans Prometheus, Grafana, DataDog, New Relic, CloudWatch, Jaeger, Zipkin, AWS X-Ray, OpenTelemetry, the ELK Stack, Loki, PagerDuty, Terraform, and enterprise ITSM/compliance tooling.

Who it's for

SRE and observability engineers designing monitoring, tracing, and alerting for enterprise-scale systems who need SLI/SLO frameworks, cost-optimized tool selection, and alert-noise-reduction strategies rather than a single dashboard setup.

FAQ

Common questions

Discussion

Questions & comments · 0

Sign In Sign in to leave a comment.