Agent Skills

alloy

devopsgrafana4.2K installs

Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo / Pyroscope. Covers the Alloy config language (blocks, `sys.env`, component refs), `prometheus.scrape` → `remote_write`, `loki.source.file` + `loki.process` → `loki.write`, `otelcol.receiver.otlp` → `otelcol.exporter.otlp`, `pyroscope.scrape`, K8s / Docker / EC2 discovery, relabeling, modules (`import.

Install

npx skills add https://github.com/grafana/skills --skill alloy
SKILL.md

Grafana Alloy

Docs: https://grafana.com/docs/alloy/latest/

OpenTelemetry-compatible collector — one binary for metrics + logs + traces + profiles.

Prerequisites

  • alloy CLI installed (brew install grafana/grafana/alloy, apt install alloy, or grafana/alloy Docker image)
  • An endpoint to ship to (Grafana Cloud, Prometheus, Loki, Tempo, Pyroscope)
  • API key + username for that endpoint, exported as GRAFANA_API_KEY etc.

Common Workflows

1. Write + validate a config locally

// config.alloy — metrics → Grafana Cloud
prometheus.scrape "app" {
  targets = [{"__address__" = "localhost:9090"}]
  forward_to = [prometheus.remote_write.cloud.receiver]
  scrape_interval = "30s"
}

prometheus.remote_write "cloud" {
  endpoint {
    url = sys.env("PROMETHEUS_URL")
    basic_auth {
      username = sys.env("PROM_USER")
      password = sys.env("GRAFANA_API_KEY")
    }
  }
}
# 1. Format + syntax-check (catches typos before run)
alloy fmt config.alloy
alloy validate config.alloy

# 2. Run it
alloy run config.alloy
# or as a service: systemctl restart alloy

# 3. Verify all components are healthy via the UI on port 12345
curl -s http://localhost:12345/api/v0/web/components \
  | jq '.[] | select(.health.state != "healthy") | {id, state:.health.state, msg:.health.message}'
# Expect: empty (nothing unhealthy). Otherwise the row shows the failing component + reason.

# 4. Verify samples are flowing
curl -s http://localhost:12345/metrics \
  | grep -E '^prometheus_remote_storage_(samples_total|enqueue_retries_total)' | head
# samples_total should be > 0 and rising; retries should be 0.

2. Add log shipping (file → Loki)

loki.source.file "app_logs" {
  targets    = [{ __path__ = "/var/log/app/*.log", job = "app" }]
  forward_to = [loki.write.cloud.receiver]
}

loki.write "cloud" {
  endpoint {
    url = sys.env("LOKI_URL")
    basic_auth {
      username = sys.env("LOKI_USER")
      password = sys.env("GRAFANA_API_KEY")
    }
  }
}
# Verify in Grafana → Explore → Loki:
#   {job="app"}
# Expect lines streaming. If empty:
#   - check /var/log/app/*.log actually exists + readable by the alloy user
#   - http://localhost:12345 → loki.source.file.app_logs → "Targets" tab

3. Receive OTLP traces and ship to Tempo

otelcol.receiver.otlp "default" {
  grpc { endpoint = "0.0.0.0:4317" }
  http { endpoint = "0.0.0.0:4318" }
  output { traces = [otelcol.exporter.otlp.tempo.input] }
}

otelcol.exporter.otlp "tempo" {
  client {
    endpoint = "tempo-xxx.grafana.net/tempo:443"
    auth     = otelcol.auth.basic.grafana_cloud.handler
  }
}

otelcol.auth.basic "grafana_cloud" {
  username = sys.env("TEMPO_USER")
  password = sys.env("GRAFANA_API_KEY")
}
# Verify in Grafana → Explore → Tempo → service.name = your-service
# At the Alloy level, watch the receiver:
curl -s http://localhost:12345/metrics | grep otelcol_receiver_accepted_spans

Full pattern set (Kubernetes discovery + relabel, complete Cloud pipeline for all 4 signals): references/collection-patterns.md.

Reference

Troubleshooting

  • alloy validate → "component not found" → version too old; alloy --version and upgrade
  • UI shows component unhealthy → click into it on http://localhost:12345 for the live error
  • prometheus_remote_storage_samples_dropped_total rising → check enqueue_retries_total and the remote_write endpoint URL + creds
  • No spans landing in Tempo → check otelcol_receiver_refused_spans and the exporter's otelcol_exporter_sent_spans / otelcol_exporter_send_failed_spans

Resources

Related skills

azure-diagnosticsmicrosoft608KDebug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, aazure-preparemicrosoft608KPrepare azd-based Azure projects for deployment: generates azure.yaml, infrastructure (Bicep/Terraform), and Dockerfiles for the Azure Developer CLI (azd) workflow. USE ONLY when the user explicitly wants to use azd as the deployment tool, or the project already has an azure.yaml file. DO NOT USE FOR: non-azd deployments, Python App Service code-only deploys (use python-appservice-deploy), or cross-cloud migration (use azure-cloud-migrate). WHEN: prepare app for azd, create azure.yaml, set up azazure-aimicrosoft608KUse for Azure AI: Search, Speech, OpenAI, Document Intelligence. Helps with search, vector/hybrid search, speech-to-text, text-to-speech, transcription, OCR. WHEN: AI Search, query search, vector search, hybrid search, semantic search, speech-to-text, text-to-speech, transcribe, OCR, convert text to speech.azure-deploymicrosoft607KExecute Azure deployments for ALREADY-PREPARED applications that have existing .azure/deployment-plan.md and infrastructure files. DO NOT use this skill when the user asks to CREATE a new application — use azure-prepare instead. This skill runs azd up, azd deploy, terraform apply, and az deployment commands with built-in error recovery. Requires .azure/deployment-plan.md from azure-prepare and validated status from azure-validate. WHEN: \"run azd up\", \"run azd deploy\", \"execute deployment\",

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers