Agent Skills

tao-run-on-brev

devopsnvidia1.6K installs

Run a TAO training/evaluation/inference container on an NVIDIA Brev GPU instance. Instance provisioning (create/search/stop/delete/login) is delegated to the official brev-cli agent skill or the Brev MCP server; this skill covers only the TAO-specific part — running the container over `brev exec` via the four-verb docker contract. Trigger phrases include "run on Brev", "Brev GPU instance", "TAO on Brev", "submit job to Brev".

Install

npx skills add https://github.com/nvidia/skills --skill tao-run-on-brev
SKILL.md

Brev — TAO execution glue

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

NVIDIA Brev provides on-demand GPU instances (pre-loaded with NVIDIA drivers, CUDA, Docker, and the NVIDIA Container Toolkit). Brev is instance-based: you provision an instance, run commands on it over brev exec, and delete it when done.

This skill is deliberately thin. Provisioning and managing instances — create, search by GPU/price, start/stop, delete, login — is owned by NVIDIA Brev's own agent skill, not duplicated here. This skill covers only the TAO-specific part: running a TAO container on a reached instance through the four-verb docker contract, deferring the container-how to tao-run-on-docker over brev exec.

Provisioning: use the official Brev skill or MCP

NVIDIA Brev publishes an agent skill that manages instances in natural language ("create an A100 instance", "search for GPUs under $3/hr", "stop all my instances"). Install it once — it self-registers into your agent's skills dir and is discovered at runtime:

curl -fsSL https://raw.githubusercontent.com/brevdev/brev-cli/main/scripts/install-agent-skill.sh | bash
# installs to ~/.claude/skills/brev-cli/ , ~/.codex/skills/brev-cli/ , ~/.agents/skills/brev-cli/

Or connect the Brev MCP server (https://docs.nvidia.com/brev/_mcp/server). Either one owns login/auth quirks, placement IDs, GPU search, and teardown flags. It does not cover container execution on the instance — that is this skill.

Preflight for this skill: the brev CLI is on PATH and logged in (headless: brev login --token "$BREV_API_TOKEN" before any other call), and you can reach a target instance — poll with a two-word command until it succeeds before issuing real work (a fresh instance reports RUNNING before sshd is up):

for i in $(seq 1 60); do brev exec <instance> "echo ok" 2>/dev/null | grep -qx ok && break; sleep 5; done
brev exec <instance> "echo ok" 2>/dev/null | grep -qx ok || { echo "instance not exec-ready"; exit 1; }

The probe must be two words, quoted as one argument. A single-token probe (brev exec <instance> -- true) passes even when every real command is broken, because brev exec [instance...] <command> treats only the LAST positional as the command — so a lone true lands in the right slot by accident while docker run ... does not. See brev exec argument form below.

Allow ≥ 600 s for the first brev exec on a new instance (SSH bring-up + first container pull); a 60–120 s wrapper timeout truncates startup and looks like a spurious exec failed.

Storage

No shared NFS/Lustre — storage tier B/C via tao-data-io: stage inputs from S3 to the instance's local disk (or fetch in-container) and upload results to S3 before deleting the instance. Instance-local ~/ persists across stop/start but not across delete/create, so the results upload must precede teardown.

Execution — the four verbs (a compound over Docker)

Brev is a compound consumer: submit reaches an instance, then defers the container-how to the four docker verbs (tao-run-on-docker) run over brev exec. It is not a symmetric peer — teardown must additionally delete the instance to stop billing. $BANK = ${TAO_SKILL_BANK_PATH}.

  • submit — reach an instance (provision/reuse via the official Brev skill or MCP; reuse an existing instance by its instance_id; wait for readiness, above). Lint the assembled command, open the record to mint $JOB_ID before launch, then run the docker submit verb inside the instance and mark RUNNING:

    redact_secrets.py lint <<<"$REMOTE_CMD"     # no inline secrets; creds as -e VAR
    JOB_ID=$("$BANK/scripts/tao_job_record.py" open \
      --platform brev --image "$IMG" \
      --network-arch "$ARCH" --action "$ACTION" \
      --storage-tier "$TIER" --results-root "$RESULTS_ROOT")
    brev exec <instance> "docker inspect '$JOB_ID' >/dev/null 2>&1 && { echo '$JOB_ID already submitted'; exit 0; }; docker run -d --name '$JOB_ID' --label 'tao-job=$JOB_ID' ..."
    "$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state RUNNING \
      --backend-ref "<instance>/$JOB_ID"       # instance is part of the ref: the
                                               # container is unreachable without it
    
  • status / logs — brev exec <instance> "docker inspect $JOB_ID" / brev exec <instance> "docker logs $JOB_ID", mapped to the vocab exactly as the docker verbs do. Recover <instance> from the record's backend-ref.

  • cancel / teardown — remove the container, then for an ephemeral instance delete it (stops billing), then mark the record. Never leave an ephemeral instance running:

    brev exec <instance> "docker rm -f $JOB_ID"
    brev delete <instance>                      # ephemeral instances only
    "$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state CANCELED --source agent
    

brev exec argument form

The remote command is one argument. The CLI signature is brev exec [instance...] <command>: every positional except the last is an instance name, and -- only ends flag parsing — it does not group the words after it. So brev exec <inst> -- docker inspect "$JOB_ID" is read as instances <inst> docker inspect plus command "$JOB_ID", and fails with could not look up instance "docker" / ssh: illegal option -- - (exit 255) — an error that reads like an instance or SSH fault but is a syntax fault. Quote every remote command as a single string, exactly as brev exec --help shows.

NGC auth once per instance — never put NGC_KEY on argv (it lands in the remote process table); pipe it to --password-stdin:

IMG=nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt  # versions-key: images.tao_toolkit.pyt

# NGC auth (one-time per instance) — value never on argv.
# Single-quoted locally so $NGC_KEY expands in the instance's shell; export it
# there first (or pipe it in from the local shell, if the instance has no copy).
brev exec <instance> 'printf %s "$NGC_KEY" | docker login nvcr.io -u "$oauthtoken" --password-stdin'

# Verify auth without reading ~/.docker/config.json. Failure before a successful
# login = not authenticated; failure after = the key's org lacks entitlement.
brev exec <instance> "docker manifest inspect $IMG >/dev/null && echo AUTH_OK || echo AUTH_FAIL"

# Pull BEFORE the GPU run. `docker run` would pull implicitly, but the instance
# bills from boot, so a multi-GB first-time TAO pull is billed GPU-idle time.
# Pulling as its own step also separates a pull failure (auth/entitlement) from
# a training failure in the logs.
brev exec <instance> "docker image inspect $IMG >/dev/null 2>&1 || docker pull $IMG"

# Run a TAO job (the docker `submit` verb, over brev exec)
brev exec <instance> "docker inspect '$JOB_ID' >/dev/null 2>&1 && { echo '$JOB_ID already submitted'; exit 0; }; docker run -d --name '$JOB_ID' --label 'tao-job=$JOB_ID' --gpus all -v ~/data:/data -e NGC_KEY '$IMG' visual_changenet train -e /data/spec.yaml"

Multi-GPU and multi-node

Multi-node is not supported on Brev — instance-based, no cross-instance coordination. Multi-GPU on a single instance is supported (up to 8× H100 / A100 / L40S); torchrun --nproc-per-node=N or PyTorch DDP work within the instance.

Related skills

azure-diagnosticsmicrosoft608KDebug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, aazure-preparemicrosoft608KPrepare azd-based Azure projects for deployment: generates azure.yaml, infrastructure (Bicep/Terraform), and Dockerfiles for the Azure Developer CLI (azd) workflow. USE ONLY when the user explicitly wants to use azd as the deployment tool, or the project already has an azure.yaml file. DO NOT USE FOR: non-azd deployments, Python App Service code-only deploys (use python-appservice-deploy), or cross-cloud migration (use azure-cloud-migrate). WHEN: prepare app for azd, create azure.yaml, set up azazure-aimicrosoft608KUse for Azure AI: Search, Speech, OpenAI, Document Intelligence. Helps with search, vector/hybrid search, speech-to-text, text-to-speech, transcription, OCR. WHEN: AI Search, query search, vector search, hybrid search, semantic search, speech-to-text, text-to-speech, transcribe, OCR, convert text to speech.azure-deploymicrosoft607KExecute Azure deployments for ALREADY-PREPARED applications that have existing .azure/deployment-plan.md and infrastructure files. DO NOT use this skill when the user asks to CREATE a new application — use azure-prepare instead. This skill runs azd up, azd deploy, terraform apply, and az deployment commands with built-in error recovery. Requires .azure/deployment-plan.md from azure-prepare and validated status from azure-validate. WHEN: \"run azd up\", \"run azd deploy\", \"execute deployment\",

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers