Agent Skills

cost-management

devopsgrafana3.5K installs

Cut your Grafana Cloud bill by attributing spend to teams and reducing telemetry volume. Covers FOCUS-compliant billing dashboards, cost-attribution labels in Alloy, Adaptive Metrics (cardinality reduction), Adaptive Logs (drop/sample), Adaptive Traces (tail sampling), usage alerts, and an optimization checklist. Use when investigating a high Grafana Cloud bill, attributing observability cost to a team or service, reducing active series / log bytes / trace spans, or setting up usage / quota aler

Install

npx skills add https://github.com/grafana/skills --skill cost-management
SKILL.md

Grafana Cloud Cost Management

Docs: https://grafana.com/docs/grafana-cloud/cost-management-and-billing/

Reduce metric / log / trace spend with Adaptive signals + cost-attribution labels.

Prerequisites

  • A Grafana Cloud stack with Adaptive Metrics / Logs / Traces enabled (visible under Cost Management)
  • Alloy (or Grafana Agent) ingesting telemetry, with API key in scope metrics:write + logs:write (+ traces:write)
  • Admin access to the stack to apply Adaptive recommendations

Common Workflows

1. Attribute cost to a team / service

# 1. Add external labels in Alloy (metrics + logs configs)
prometheus.remote_write "cloud" {
  endpoint { url = sys.env("PROMETHEUS_URL") /* ... */ }
  external_labels = { team = "platform", project = "checkout-service" }
}
# 2. Reload Alloy
curl -X POST http://localhost:12345/-/reload

# 3. Verify labels arrived in Grafana Cloud
#    In Explore, run:  count by (team, project) ({__name__=~".+"})
#    Then visit Cost Management → group by `team` / `project`

See references/adaptive-signals.md for the full Alloy snippet.

2. Cut metric cardinality with Adaptive Metrics

# 1. Pull recommendations
curl https://<stack>.grafana.net/api/plugins/grafana-adaptive-metrics-app/resources/v1/recommendations \
  -H "Authorization: Bearer <token>" | jq '.recommendations | length'

# 2. In the UI: Grafana Cloud → Adaptive Metrics → review rules sorted by series-reduction impact
# 3. Test in "Preview" mode before applying
# 4. Apply (takes effect within 5 min)

# 5. Verify — series count should drop on the affected metrics
#    Before applying, capture baseline:
#      count({__name__="http_request_duration_seconds_bucket"})
#    Wait 10 min after apply, run again — expect 10x+ reduction for high-card metrics.

# Rollback if needed: open the rule in the UI → Disable, or DELETE /v1/rules/<id>.

3. Drop noisy logs in Alloy

# 1. Add a filter stage (see references/adaptive-signals.md for the full block)
loki.process "filter_logs" {
  forward_to = [loki.write.cloud.receiver]
  stage.drop { expression = ".*GET /health.*" }
}
# 2. Reload Alloy
curl -X POST http://localhost:12345/-/reload

# 3. Verify the filter — health logs should NOT appear in Logs Drilldown
#    LogQL check (should return 0):
#      sum(rate({app="my-app"} |= "GET /health" [5m]))
#    Bytes-ingested should also drop. Compare 24h before/after:
#      sum(increase(loki_ingester_chunk_size_bytes_sum[24h])) by (namespace)

4. Set a usage alert before you hit quota

See references/alerts-and-queries.md for ready-to-paste rules (MetricsUsageHigh, LogsIngestionHigh).

Optimization checklist

  • Apply Adaptive Metrics recommendations — typically reduces series 40-60%
  • Drop health/readiness probe logs in Alloy
  • Tail-sample traces to 5-10% + keep errors / slow spans
  • Add team + project external labels to every Alloy config
  • Set usage alerts at 80% of quota
  • Replace expensive ad-hoc queries with recording rules

References

Resources

Related skills

azure-diagnosticsmicrosoft608KDebug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, aazure-preparemicrosoft608KPrepare azd-based Azure projects for deployment: generates azure.yaml, infrastructure (Bicep/Terraform), and Dockerfiles for the Azure Developer CLI (azd) workflow. USE ONLY when the user explicitly wants to use azd as the deployment tool, or the project already has an azure.yaml file. DO NOT USE FOR: non-azd deployments, Python App Service code-only deploys (use python-appservice-deploy), or cross-cloud migration (use azure-cloud-migrate). WHEN: prepare app for azd, create azure.yaml, set up azazure-aimicrosoft608KUse for Azure AI: Search, Speech, OpenAI, Document Intelligence. Helps with search, vector/hybrid search, speech-to-text, text-to-speech, transcription, OCR. WHEN: AI Search, query search, vector search, hybrid search, semantic search, speech-to-text, text-to-speech, transcribe, OCR, convert text to speech.azure-deploymicrosoft607KExecute Azure deployments for ALREADY-PREPARED applications that have existing .azure/deployment-plan.md and infrastructure files. DO NOT use this skill when the user asks to CREATE a new application — use azure-prepare instead. This skill runs azd up, azd deploy, terraform apply, and az deployment commands with built-in error recovery. Requires .azure/deployment-plan.md from azure-prepare and validated status from azure-validate. WHEN: \"run azd up\", \"run azd deploy\", \"execute deployment\",

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers