omni-inference
The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents.
Install
npx skills add https://github.com/diegosouzapw/omniroute --skill omni-inferenceOverview
The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents.
Authentication
All requests require a valid Bearer token or session cookie. Obtain a token via POST /api/auth/login or configure REQUIRE_API_KEY=false for local development.
Endpoints
POST /api/v1/session-leasesGET /api/v1/searchPOST /api/v1/searchPOST /api/v1/chat/completionsGET /api/v1/wsPOST /api/v1/providers/{provider}/chat/completionsPOST /api/v1/api/chatPOST /api/v1/messagesPOST /api/v1/messages/count_tokensPOST /api/v1/responsesPOST /api/v1/embeddingsGET /api/v1/multimodal-embeddingsPOST /api/v1/multimodal-embeddingsPOST /api/v1/providers/{provider}/embeddingsPOST /api/v1/images/generationsPOST /api/v1/providers/{provider}/images/generationsPOST /api/v1/audio/speechPOST /api/v1/audio/transcriptionsPOST /api/v1/moderationsPOST /api/v1/rerankGET /api/v1GET /api/v1/providers/{provider}/modelsGET /api/v1/management/proxy-subscriptionsPOST /api/v1/management/proxy-subscriptionsGET /api/v1/management/proxy-subscriptions/{id}PATCH /api/v1/management/proxy-subscriptions/{id}DELETE /api/v1/management/proxy-subscriptions/{id}GET /api/v1/management/proxy-subscriptions/{id}/nodesPOST /api/v1/management/proxy-subscriptions/{id}/refreshPOST /api/v1/ocrPOST /api/v1/audio/translationsGET /api/v1/voicesPOST /api/v1/speech-to-textPOST /api/v1/text-to-speech/{voiceId}GET /api/v1/explain/routingGET /api/v1/providers/suggested-modelsGET /api/v1/provider-plugin-manifestGET /api/v1/{omnirouteCatchAll}POST /api/v1/{omnirouteCatchAll}PUT /api/v1/{omnirouteCatchAll}PATCH /api/v1/{omnirouteCatchAll}DELETE /api/v1/{omnirouteCatchAll}GET /api/v1/accounts/{id}/limitsPUT /api/v1/accounts/{id}/limitsGET /api/v1/agents/credentialsPOST /api/v1/agents/credentialsGET /api/v1/agents/healthGET /api/v1/agents/tasksPOST /api/v1/agents/tasksDELETE /api/v1/agents/tasksGET /api/v1/agents/tasks/{id}POST /api/v1/agents/tasks/{id}DELETE /api/v1/agents/tasks/{id}POST /api/v1/antigravityGET /api/v1/auto-combo/{channel}/candidatesGET /api/v1/batchesPOST /api/v1/batchesGET /api/v1/batches/{id}DELETE /api/v1/batches/{id}POST /api/v1/batches/{id}/cancelDELETE /api/v1/batches/delete-completedPOST /api/v1/classifyGET /api/v1/combosPOST /api/v1/completionsGET /api/v1/filesPOST /api/v1/filesGET /api/v1/files/{id}DELETE /api/v1/files/{id}GET /api/v1/files/{id}/contentPOST /api/v1/images/editsGET /api/v1/images/upscalePOST /api/v1/images/upscalePOST /api/v1/issues/reportGET /api/v1/management/proxiesPOST /api/v1/management/proxiesPATCH /api/v1/management/proxiesDELETE /api/v1/management/proxiesGET /api/v1/management/proxies/assignmentsPUT /api/v1/management/proxies/assignmentsPUT /api/v1/management/proxies/bulk-assignGET /api/v1/management/proxies/healthGET /api/v1/me/statusGET /api/v1/muse-code/modelsGET /api/v1/music/generationsPOST /api/v1/music/generationsGET /api/v1/providers/{provider}/limitsPUT /api/v1/providers/{provider}/limitsGET /api/v1/quotas/checkGET /api/v1/registered-keysPOST /api/v1/registered-keysGET /api/v1/registered-keys/{id}DELETE /api/v1/registered-keys/{id}POST /api/v1/registered-keys/{id}/revokePOST /api/v1/relay/chat/completionsPOST /api/v1/relay/chat/completions/bifrostPOST /api/v1/responses/{path}GET /api/v1/search/analyticsPOST /api/v1/segmentGET /api/v1/video-bridge/drilldownDELETE /api/v1/video-bridge/drilldownGET /api/v1/videos/generationsPOST /api/v1/videos/generationsGET /api/v1/vscode/{token}POST /api/v1/vscode/{token}/api/chatPOST /api/v1/vscode/{token}/api/showGET /api/v1/vscode/{token}/api/tagsGET /api/v1/vscode/{token}/api/versionPOST /api/v1/vscode/{token}/chat/completionsGET /api/v1/vscode/{token}/combosGET /api/v1/vscode/{token}/modelsPOST /api/v1/vscode/{token}/responsesPOST /api/v1/vscode/{token}/v1/chat/completionsGET /api/v1/vscode/{token}/v1/modelsGET /api/v1/vscode/combos/{token}/{{slug}}POST /api/v1/vscode/combos/{token}/{{slug}}GET /api/v1/vscode/raw/{token}POST /api/v1/vscode/raw/{token}/api/chatPOST /api/v1/vscode/raw/{token}/api/showGET /api/v1/vscode/raw/{token}/api/tagsGET /api/v1/vscode/raw/{token}/api/versionPOST /api/v1/vscode/raw/{token}/chat/completionsGET /api/v1/vscode/raw/{token}/combosGET /api/v1/vscode/raw/{token}/modelsPOST /api/v1/vscode/raw/{token}/responsesPOST /api/v1/vscode/raw/{token}/v1/chat/completionsGET /api/v1/vscode/raw/{token}/v1/modelsPOST /api/v1/web/fetch
Payloads
See the full OpenAPI specification at GET /api/openapi/spec or docs/openapi.yaml for detailed request/response schemas.
Chat completions
Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.
Endpoints
POST $OMNIROUTE_URL/v1/chat/completions— OpenAI formatPOST $OMNIROUTE_URL/v1/messages— Anthropic Messages formatPOST $OMNIROUTE_URL/v1/responses— OpenAI Responses API
Discover
curl $OMNIROUTE_URL/v1/models | jq '.data[].id'
Combos (e.g. auto, cost-optimized, subscription) auto-fallback through multiple providers.
OpenAI format example
curl -X POST $OMNIROUTE_URL/v1/chat/completions \
-H "Authorization: Bearer $OMNIROUTE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-7",
"messages": [{"role": "user", "content": "Refactor this function"}],
"stream": true
}'
Anthropic format example
curl -X POST $OMNIROUTE_URL/v1/messages \
-H "Authorization: Bearer $OMNIROUTE_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-7",
"max_tokens": 4096,
"messages": [{"role": "user", "content": "Hi"}]
}'
Tool use
Supports OpenAI tools array and Anthropic tools block. Tool results
auto-compressed via RTK (47 filters: git-diff, grep, test-jest, terraform-plan,
docker-logs, etc.) — 20-40% token savings. Disable per-request with
X-Omniroute-Rtk: off header.
Reasoning / thinking
Anthropic extended thinking and OpenAI Responses reasoning blocks are forwarded verbatim. Cached automatically via reasoning cache.
Errors
401→ invalid API key400 invalid_model→ model not in registry; check/v1/models503 circuit_open→ provider circuit breaker tripped; retry later or use combo429 rate_limited→ honorRetry-After; consider using a combo for auto-fallback
Image generation
Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.
Endpoints
POST $OMNIROUTE_URL/v1/images/generations— Text-to-imagePOST $OMNIROUTE_URL/v1/images/edits— Image edit (mask)POST $OMNIROUTE_URL/v1/images/variations— Variations
Discover
curl $OMNIROUTE_URL/v1/models/image | jq '.data[]'
Returns { id, owned_by, sizes:[...], capabilities:[...] } per model.
Generate example
curl -X POST $OMNIROUTE_URL/v1/images/generations \
-H "Authorization: Bearer $OMNIROUTE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "dall-e-3",
"prompt": "a red bicycle on a wet street, photoreal",
"n": 1,
"size": "1024x1024",
"response_format": "b64_json"
}'
Response: { created, data: [{ url? or b64_json, revised_prompt }] }
Errors
400 invalid_size→ not supported by this model; check/v1/models/image400 content_policy_violation→ blocked by provider safety503→ provider unavailable; try another model in/v1/models/image
Text-to-speech
Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.
Endpoint
POST $OMNIROUTE_URL/v1/audio/speech— returns binary audio (mp3/opus/wav/flac)
Discover
curl $OMNIROUTE_URL/v1/models/tts | jq '.data[]'
Each entry includes voices:[...] for the available voice names per provider.
Example
curl -X POST $OMNIROUTE_URL/v1/audio/speech \
-H "Authorization: Bearer $OMNIROUTE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "Hello from OmniRoute.",
"voice": "alloy",
"response_format": "mp3"
}' --output speech.mp3
Voices
Voice names vary by provider. Check /v1/models/tts — each entry has voices:[...].
Common OpenAI voices: alloy, echo, fable, onyx, nova, shimmer.
Errors
400 invalid_voice→ voice not supported by this model400 input_too_long→ input exceeds model character limit503→ provider unavailable; try another model in/v1/models/tts
Speech-to-text
Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.
Endpoints
POST $OMNIROUTE_URL/v1/audio/transcriptions— multipart upload, returns textPOST $OMNIROUTE_URL/v1/audio/translations— transcribe + translate to English
Discover
curl $OMNIROUTE_URL/v1/models/stt | jq '.data[]'
Example
curl -X POST $OMNIROUTE_URL/v1/audio/transcriptions \
-H "Authorization: Bearer $OMNIROUTE_KEY" \
-F "file=@audio.mp3" \
-F "model=whisper-1" \
-F "response_format=verbose_json"
Response: { text, language, duration, segments?:[{ start, end, text }] }
Supported formats
Audio: mp3, mp4, mpeg, mpga, m4a, wav, webm.
Response formats: json, text, srt, verbose_json, vtt.
Errors
400 invalid_file_format→ unsupported audio format400 file_too_large→ exceeds provider limit (usually 25MB)503→ provider unavailable; try another model in/v1/models/stt
Embeddings
Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.
Endpoint
POST $OMNIROUTE_URL/v1/embeddings
Discover
curl $OMNIROUTE_URL/v1/models/embedding | jq '.data[]'
Each entry: { id, owned_by, dimensions, max_input_tokens }.
Example
curl -X POST $OMNIROUTE_URL/v1/embeddings \
-H "Authorization: Bearer $OMNIROUTE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-large",
"input": ["first text", "second text"],
"encoding_format": "float"
}'
Response: { data:[{ embedding:[...], index }], usage:{ prompt_tokens, total_tokens } }
Batch input
input accepts a string or array of strings (up to provider batch limit, typically 2048 items).
Errors
400 input_too_long→ input exceedsmax_input_tokensfor this model400 invalid_encoding_format→ usefloatorbase64503→ provider unavailable; try another model in/v1/models/embedding
Web search
Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.
Endpoint
POST $OMNIROUTE_URL/v1/web/search— unified search format
Discover
curl $OMNIROUTE_URL/v1/models/web | jq '.data[] | select(.kind == "webSearch")'
Example
curl -X POST $OMNIROUTE_URL/v1/web/search \
-H "Authorization: Bearer $OMNIROUTE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tavily/search",
"query": "OmniRoute github latest release",
"max_results": 5,
"include_answer": true
}'
Response: { answer?, results:[{ url, title, content, score }] }
Parameters
| Field | Type | Description |
|---|---|---|
model |
string | Provider model from /v1/models/web |
query |
string | Search query |
max_results |
number | Max results (default: 5) |
include_answer |
boolean | Include AI-synthesized answer |
search_depth |
string | basic or advanced (Tavily) |
Errors
400 query_too_long→ shorten the search query503→ provider unavailable; try another model in/v1/models/web
Web fetch
Requires OMNIROUTE_URL and OMNIROUTE_KEY. See entry-point SKILL for setup.
Endpoint
POST $OMNIROUTE_URL/v1/web/fetch
Discover
curl $OMNIROUTE_URL/v1/models/web | jq '.data[] | select(.kind == "webFetch")'
Example
curl -X POST $OMNIROUTE_URL/v1/web/fetch \
-H "Authorization: Bearer $OMNIROUTE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jina/reader",
"url": "https://anthropic.com",
"format": "markdown"
}'
Response: { url, title, markdown, links?:[...], images?:[...] }
Parameters
| Field | Type | Description |
|---|---|---|
model |
string | Provider from /v1/models/web (e.g. jina/reader, firecrawl/scrape) |
url |
string | URL to fetch |
format |
string | markdown (default), html, text |
Errors
400 invalid_url→ URL must be http/https403 blocked→ provider blocked by target site; try a different model503→ provider unavailable; try another model in/v1/models/web
