Testing & Quality Checks

Observability and Incident Diagnostics Designer

Design logs, metrics, traces, alerts, dashboards, SLOs, and runbooks that support fast diagnosis and reliable incident response.

Last updated Jul 11, 2026
FreeClaudeChatGPTCursor
TL;DR

Observability and Incident Diagnostics Designer is a free AI skill for testing & quality checks. Design logs, metrics, traces, alerts, dashboards, SLOs, and runbooks that support fast diagnosis and reliable incident response. It works with Claude, ChatGPT, Cursor and is ready to use out of the box.

Download Skill.md Package

About this skill

Observability and Incident Diagnostics Designer helps teams understand distributed systems in production. It defines service-level indicators, structured logs, metrics, traces, correlation IDs, dashboards, alerts, ownership, retention, diagnostic workflows, and incident runbooks.

What it does

The skill analyzes system architecture, critical user journeys, failure modes, dependencies, and operational responsibilities; defines the minimum useful telemetry; creates alert and dashboard standards; maps symptoms to diagnostic evidence; and produces a phased observability implementation plan.

What is included

  • Critical journey and failure map
  • SLI and SLO definitions
  • Logging specification
  • Metrics and tracing plan
  • Alerting rules
  • Dashboard design
  • Diagnostic runbooks
  • Implementation roadmap

How to use it

1. Download the observability-and-incident-diagnostics-designer-SKILL.md file
2. Upload it to your AI or platform workspace
3. Describe the architecture, critical journeys, incidents, and current tooling
4. Add availability, latency, retention, and compliance requirements
5. Use the design to implement and improve operational visibility

Examples

Example input
Design observability for a SaaS platform with a web app, API, PostgreSQL database, background workers, Redis, object storage, and third-party payment integration.
Example output
A complete observability design with SLIs, SLOs, logs, traces, metrics, correlation IDs, dashboards, payment and job alerts, diagnostic runbooks, ownership, and rollout priorities.

FAQ

What is this skill for?
It designs an observability and diagnostic system for production software.
Does it include SLOs?
Yes. It defines service-level indicators, objectives, error budgets, and ownership where appropriate.
Can it reduce alert fatigue?
Yes. It focuses alerts on actionable symptoms, user impact, thresholds, routing, and deduplication.
Does it require distributed tracing?
Not always. It recommends tracing where request or job paths cross important service boundaries.
Can it create incident runbooks?
Yes. It maps symptoms to queries, dashboards, checks, mitigations, escalation, and recovery verification.
How is this different from installing monitoring tools?
It defines what to measure, why it matters, how to diagnose, and who responds.

Related Skills

Testing & Quality ChecksFree

Synthetic Monitoring Test Designer

Design synthetic monitoring for critical journeys with checks, locations, test data, thresholds, alerts, diagnosis, and maintenance.

ClaudeChatGPT
#synthetic monitoring#production testing#reliability checks
Testing & Quality ChecksFree

API Contract Testing Designer

Design API contract tests for providers and consumers using schemas, examples, compatibility, mocks, CI gates, and failure cases.

ClaudeChatGPTCursor
#API contract testing#consumer driven contracts#Pact
Testing & Quality ChecksFree

Automated Test Suite Architect

Design maintainable automated test suites with test layers, boundaries, fixtures, mocks, coverage goals, CI execution, and ownership.

ClaudeChatGPTCursor
#automated testing#test architecture#unit testing

Related Prompts

Free

Automation Failure Monitoring & Recovery

Create an observability and incident-recovery design for business automations, including logs, alerts, retry policy, ownership, replay, and post-incident review.

ClaudeChatGPT
#automation monitoring#incident response#workflow reliability
Free

Agent Observability Specification Builder

Specify the traces, metrics, logs, alerts, and review views needed to understand an agent's decisions and diagnose failures.

ClaudeChatGPT
#agent-observability#ai-monitoring#distributed-tracing
Free

Production Code-Path Reconstructor

Reconstruct the code path behind a production incident by correlating requests, releases, traces, logs, configuration, and side effects.

ClaudeChatGPT
#production-debugging#incident-analysis#distributed-tracing