DevOps
1,208 skills.
Browse
monitoring-observabilityahmedasmarMonitoring and observability strategy, implementation, and troubleshooting. Use this skill whenever the user mentions monitoring, observability, metrics, logs, traces, alerting, SLOs, Prometheus, Grafana, Datadog, Loki, or OpenTelemetry. Triggers include designing metrics strategy (Four Golden Signals, RED/USE), setting up Prometheus/Grafana/Loki, creating alerts or dashboards, calculating SLOs and error budgets, instrumenting with OpenTelemetry, analyzing performance issues, choosing between mocicd-pipeline-generatorailabs-393This skill should be used when creating or configuring CI/CD pipeline files for automated testing, building, and deployment. Use this for generating GitHub Actions workflows, GitLab CI configs, CircleCI configs, or other CI/CD platform configurations. Ideal for setting up automated pipelines for Node.js/Next.js applications, including linting, testing, building, and deploying to platforms like Vercel, Netlify, or AWS.docker-containerizationailabs-393This skill should be used when containerizing applications with Docker, creating Dockerfiles, docker-compose configurations, or deploying containers to various platforms. Ideal for Next.js, React, Node.js applications requiring containerization for development, production, or CI/CD pipelines. Use this skill when users need Docker configurations, multi-stage builds, container orchestration, or deployment to Kubernetes, ECS, Cloud Run, etc.alert-managementaj-geddesImplement comprehensive alert management with PagerDuty, escalation policies, and incident coordination. Use when setting up alerting systems, managing on-call schedules, or coordinating incident response.api-gateway-configurationaj-geddesConfigure API gateways for routing, authentication, rate limiting, and request/response transformation. Use when deploying microservices, setting up reverse proxies, or managing API traffic.app-store-deploymentaj-geddesDeploy iOS and Android apps to App Store and Google Play. Covers signing, versioning, build configuration, submission process, and release management.application-loggingaj-geddesImplement structured logging across applications with log aggregation and centralized analysis. Use when setting up application logging, implementing ELK stack, or analyzing application behavior.artifact-managementaj-geddesManage build artifacts, Docker images, and package registries. Configure artifact repositories, versioning, and distribution strategies.autoscaling-configurationaj-geddesConfigure autoscaling for Kubernetes, VMs, and serverless workloads based on metrics, schedules, and custom indicators.aws-cloudfront-cdnaj-geddesDistribute content globally using CloudFront with caching, security headers, WAF integration, and origin configuration. Use for low-latency content delivery.aws-ec2-setupaj-geddesLaunch and configure EC2 instances with security groups, IAM roles, key pairs, AMIs, and auto-scaling. Use for virtual servers and managed infrastructure.aws-lambda-functionsaj-geddesCreate and deploy serverless functions using AWS Lambda with event sources, permissions, layers, and environment configuration. Use for event-driven computing without managing servers.aws-s3-managementaj-geddesManage S3 buckets with versioning, encryption, access control, lifecycle policies, and replication. Use for object storage, static sites, and data lakes.azure-app-serviceaj-geddesDeploy and manage web apps using Azure App Service with auto-scaling, deployment slots, SSL/TLS, and monitoring. Use for hosting web applications on Azure.azure-functionsaj-geddesCreate serverless functions on Azure with triggers, bindings, authentication, and monitoring. Use for event-driven computing without managing infrastructure.backup-disaster-recoveryaj-geddesImplement backup strategies, disaster recovery plans, and data restoration procedures for protecting critical infrastructure and data.blue-green-deploymentaj-geddesImplement blue-green deployment strategies for zero-downtime releases with instant rollback capability and traffic switching between environments.canary-deploymentaj-geddesImplement canary deployment strategies to gradually roll out new versions to subset of users with automatic rollback based on metrics.cicd-pipeline-setupaj-geddesDesign and implement CI/CD pipelines with GitHub Actions, GitLab CI, Jenkins, or CircleCI. Use for automated testing, building, and deployment workflows.cloud-cost-managementaj-geddesOptimize and manage cloud costs across AWS, Azure, and GCP using reserved instances, spot pricing, and cost monitoring tools.cloud-migration-planningaj-geddesPlan and execute cloud migrations with assessment, database migration, application refactoring, and cutover strategies across AWS, Azure, and GCP.cloud-storage-optimizationaj-geddesOptimize cloud storage across AWS S3, Azure Blob, and GCP Cloud Storage with compression, partitioning, lifecycle policies, and cost management.configuration-managementaj-geddesManage application configuration including environment variables, settings management, configuration hierarchies, secret management, feature flags, and 12-factor app principles. Use for config, environment setup, or settings management.container-debuggingaj-geddesDebug Docker containers and containerized applications. Diagnose deployment issues, container lifecycle problems, and resource constraints.container-registry-managementaj-geddesManage container registries (Docker Hub, ECR, GCR) with image scanning, retention policies, and access control.correlation-tracingaj-geddesImplement distributed tracing with correlation IDs, trace propagation, and span tracking across microservices. Use when debugging distributed systems, monitoring request flows, or implementing observability.deployment-automationaj-geddesAutomate deployments across environments using Helm, Terraform, and ArgoCD. Implement blue-green deployments, canary releases, and rollback strategies.disaster-recovery-testingaj-geddesExecute comprehensive disaster recovery tests, validate recovery procedures, and document lessons learned from DR exercises.distributed-tracingaj-geddesImplement distributed tracing with Jaeger and Zipkin for tracking requests across microservices. Use when debugging distributed systems, tracking request flows, or analyzing service performance.dns-managementaj-geddesManage DNS records, routing policies, and failover configurations for high availability and disaster recovery.docker-containerizationaj-geddesCreate optimized Docker containers with multi-stage builds, security best practices, and minimal image sizes. Use when containerizing applications, creating Dockerfiles, optimizing container images, or setting up Docker Compose services.error-trackingaj-geddesImplement error tracking with Sentry for automatic exception monitoring, release tracking, and performance issues. Use when setting up error monitoring, tracking bugs in production, or analyzing application stability.feature-flag-systemaj-geddesImplement feature flags (toggles) for controlled feature rollouts, A/B testing, canary deployments, and kill switches. Use when deploying new features gradually, testing in production, or managing feature lifecycles.gcp-cloud-functionsaj-geddesDeploy serverless functions on Google Cloud Platform with triggers, IAM roles, environment variables, and monitoring. Use for event-driven computing on GCP.gcp-cloud-runaj-geddesDeploy containerized applications on Google Cloud Run with automatic scaling, traffic management, and service mesh integration. Use for container-based serverless computing.github-actions-workflowaj-geddesBuild comprehensive GitHub Actions workflows for CI/CD, testing, security, and deployment. Master workflows, jobs, steps, and conditional execution.gitlab-cicd-pipelineaj-geddesDesign and implement GitLab CI/CD pipelines with stages, jobs, artifacts, and caching. Configure runners, Docker integration, and deployment strategies.grafana-dashboardaj-geddesCreate professional Grafana dashboards with visualizations, templating, and alerts. Use when building monitoring dashboards, creating data visualizations, or setting up operational insights.health-check-endpointsaj-geddesImplement comprehensive health check endpoints for liveness, readiness, and dependency monitoring. Use when deploying to Kubernetes, implementing load balancer health checks, or monitoring service availability.infrastructure-cost-optimizationaj-geddesOptimize cloud infrastructure costs through resource rightsizing, reserved instances, spot instances, and waste reduction strategies.infrastructure-monitoringaj-geddesSet up comprehensive infrastructure monitoring with Prometheus, Grafana, and alerting systems for metrics, health checks, and performance tracking.kubernetes-deploymentaj-geddesDeploy, manage, and scale containerized applications on Kubernetes clusters with best practices for production workloads, resource management, and rolling updates.load-balancer-setupaj-geddesConfigure and deploy load balancers (HAProxy, AWS ELB/ALB/NLB) for distributing traffic, session management, and high availability.log-aggregationaj-geddesImplement centralized logging with ELK Stack, Loki, or Splunk for log collection, parsing, storage, and analysis across infrastructure.log-analysisaj-geddesAnalyze application and system logs to identify errors, patterns, and root causes. Use log aggregation tools and structured logging for effective debugging.logging-best-practicesaj-geddesImplement structured logging with JSON formats, log levels (DEBUG, INFO, WARN, ERROR), contextual logging, PII handling, and centralized logging. Use for logging, observability, log levels, structured logs, or debugging.ML Pipeline Automationaj-geddesBuild end-to-end ML pipelines with automated data processing, training, validation, and deployment using Airflow, Kubeflow, and JenkinsModel Deploymentaj-geddesDeploy machine learning models to production using Flask, FastAPI, Docker, cloud platforms (AWS, GCP, Azure), and model serving frameworksModel Monitoringaj-geddesMonitor model performance, detect data drift, concept drift, and anomalies in production using Prometheus, Grafana, and MLflowmulti-cloud-strategyaj-geddesDesign and implement multi-cloud strategies spanning AWS, Azure, and GCP with vendor lock-in avoidance, hybrid deployments, and federation.
