alert-investigation
Investigates a triggered observability alert and returns a structured diagnosis with likely cause, scope, and next steps.
Install
npx skills add https://github.com/launchdarkly/ai-tooling --skill alert-investigationSKILL.md
Alert investigation
You are investigating a specific triggered alert. Alerts arrive with structured context — an alert ID, name, threshold, value that crossed it, and a time range. Your job is to explain why it fired, assess scope, and recommend action.
Prerequisites
This skill uses the following LaunchDarkly observability MCP tools:
query-logs— query log recordsquery-traces— query distributed tracesquery-error-groups— query error groupsquery-sessions— query sessionsquery-aggregations— query aggregated/time-bucketed metricsget-keys— discover available attribute keys before filtering
Workflow
- Parse the alert context. The first turn of the conversation carries alert variables:
alertID,alertName,alertValue,group,groupValue,query,thresholdWindow,timeRange, plus a product-specific link. Use these, don't re-derive them. - Load the per-product companion. Based on the alert's product type, load the matching companion:
logs.md,traces.md,errors.md,sessions.md, ormetrics.md. Each captures the per-product investigation shape. - Run the investigation using the methodology from the investigate skill (cross-reference logs/traces/errors/sessions/metrics; cite identifiers; aggregate before paginating). Scoped to the alert's time range and filter.
- Produce a structured diagnosis. See output template below.
Output template
Alert investigations have a consistent structure so consumers (notification channels, dashboards) can parse them.
## What triggered
<1-2 sentences naming the alert, the threshold, and the value that crossed it.>
## Likely cause
<Root-cause narrative citing specific evidence: trace IDs, log timestamps, error group IDs, flag keys, deploy timing.>
## Scope
<Who or what is affected. Number of users, services, sessions, error groups. Time window of impact.>
## Next steps
<1-3 concrete actions the on-call or owner should take. Prefer specifics: "roll back flag X in env Y", "restart service Z", "investigate trace <id> for the downstream failure". Avoid "investigate further" — if you don't have a root cause, say what specifically should be investigated and how.>
When to load which companion
logs.md— log alert, log pattern alerttraces.md— latency alert, trace-error-rate alert, span-specific alerterrors.md— error-rate alert, new-error-group alert, crash-rate alertsessions.md— session-health alert, user-facing-error-rate alertmetrics.md— custom metric threshold, aggregated metric alert, composite alert
If the alert crosses product boundaries (e.g. a metric alert driven by error data), load both companions.
Guidelines
- Stay tight. Alert investigations feed notifications — keep the output structured and scannable. No preamble ("Here is my analysis..."), no repeated framing.
- Cite identifiers. Every claim in the diagnosis should reference a specific trace ID, error group ID, session ID, or log timestamp.
- If the alert appears to be noise, say so explicitly — "This alert fired because of , but the underlying behavior is within normal variance because ". Noise is a legitimate outcome; don't invent root causes.
- Don't redo the investigation you just did. The diagnosis output should let the on-call act without re-querying.
Related skills
azure-diagnosticsmicrosoft608KDebug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, aazure-preparemicrosoft608KPrepare azd-based Azure projects for deployment: generates azure.yaml, infrastructure (Bicep/Terraform), and Dockerfiles for the Azure Developer CLI (azd) workflow. USE ONLY when the user explicitly wants to use azd as the deployment tool, or the project already has an azure.yaml file. DO NOT USE FOR: non-azd deployments, Python App Service code-only deploys (use python-appservice-deploy), or cross-cloud migration (use azure-cloud-migrate). WHEN: prepare app for azd, create azure.yaml, set up azazure-aimicrosoft608KUse for Azure AI: Search, Speech, OpenAI, Document Intelligence. Helps with search, vector/hybrid search, speech-to-text, text-to-speech, transcription, OCR. WHEN: AI Search, query search, vector search, hybrid search, semantic search, speech-to-text, text-to-speech, transcribe, OCR, convert text to speech.azure-deploymicrosoft607KExecute Azure deployments for ALREADY-PREPARED applications that have existing .azure/deployment-plan.md and infrastructure files. DO NOT use this skill when the user asks to CREATE a new application — use azure-prepare instead. This skill runs azd up, azd deploy, terraform apply, and az deployment commands with built-in error recovery. Requires .azure/deployment-plan.md from azure-prepare and validated status from azure-validate. WHEN: \"run azd up\", \"run azd deploy\", \"execute deployment\",
