Observability and Incident Diagnostics Designer
Design logs, metrics, traces, alerts, dashboards, SLOs, and runbooks that support fast diagnosis and reliable incident response.
Observability and Incident Diagnostics Designer is a free AI skill for testing & quality checks. Design logs, metrics, traces, alerts, dashboards, SLOs, and runbooks that support fast diagnosis and reliable incident response. It works with Claude, ChatGPT, Cursor and is ready to use out of the box.
About this skill
Observability and Incident Diagnostics Designer helps teams understand distributed systems in production. It defines service-level indicators, structured logs, metrics, traces, correlation IDs, dashboards, alerts, ownership, retention, diagnostic workflows, and incident runbooks.
What it does
The skill analyzes system architecture, critical user journeys, failure modes, dependencies, and operational responsibilities; defines the minimum useful telemetry; creates alert and dashboard standards; maps symptoms to diagnostic evidence; and produces a phased observability implementation plan.
What is included
- Critical journey and failure map
- SLI and SLO definitions
- Logging specification
- Metrics and tracing plan
- Alerting rules
- Dashboard design
- Diagnostic runbooks
- Implementation roadmap
How to use it
1. Download the observability-and-incident-diagnostics-designer-SKILL.md file 2. Upload it to your AI or platform workspace 3. Describe the architecture, critical journeys, incidents, and current tooling 4. Add availability, latency, retention, and compliance requirements 5. Use the design to implement and improve operational visibility
Examples
Design observability for a SaaS platform with a web app, API, PostgreSQL database, background workers, Redis, object storage, and third-party payment integration.
A complete observability design with SLIs, SLOs, logs, traces, metrics, correlation IDs, dashboards, payment and job alerts, diagnostic runbooks, ownership, and rollout priorities.
FAQ
What is this skill for?
Does it include SLOs?
Can it reduce alert fatigue?
Does it require distributed tracing?
Can it create incident runbooks?
How is this different from installing monitoring tools?
Related Skills
Synthetic Monitoring Test Designer
Design synthetic monitoring for critical journeys with checks, locations, test data, thresholds, alerts, diagnosis, and maintenance.
API Contract Testing Designer
Design API contract tests for providers and consumers using schemas, examples, compatibility, mocks, CI gates, and failure cases.
Automated Test Suite Architect
Design maintainable automated test suites with test layers, boundaries, fixtures, mocks, coverage goals, CI execution, and ownership.
Related Prompts
Automation Failure Monitoring & Recovery
Create an observability and incident-recovery design for business automations, including logs, alerts, retry policy, ownership, replay, and post-incident review.
Agent Observability Specification Builder
Specify the traces, metrics, logs, alerts, and review views needed to understand an agent's decisions and diagnose failures.
Production Code-Path Reconstructor
Reconstruct the code path behind a production incident by correlating requests, releases, traces, logs, configuration, and side effects.