InstaRoute documentation

Everything InstaRoute can do and how to use it: from your first request to routing rules, budgets, savings reports, alerts and team access.

What it is

InstaRoute is an AI inference gateway and LLM router. Your apps send requests to one endpoint that works like OpenAI's or Anthropic's API. For each request, InstaRoute picks the best model from the providers you've connected, based on cost, speed and quality. It sends the request there, switches to another model if that one fails, and records exactly how much you saved.

One API, every provider

OpenAI Chat Completions, Anthropic Messages, OpenAI Responses and Embeddings formats, all on the same key.

Bring your own keys (BYOK)

You connect your own provider keys and pay providers directly at list price. Tokens are never marked up.

Smart routing

Set model to "auto" and each request goes to the cheapest model that can do the job well. Hard prompts get stronger models.

Proven savings

Every request records what it actually cost and what the most expensive suitable model would have cost.

Automatic failover

If a provider errors or is down, the request moves to the next best model, up to 4 attempts.

Private by default

Prompts and responses are never stored. Logs keep only metadata: model, tokens, cost, latency.

Quick start (5 steps)

  1. 1

    Create your account

    Register with your email. A workspace is created for you with a 7-day free trial of the Developer plan, with no card required. When the trial ends it moves to Free unless you choose a plan.
  2. 2

    Connect at least one provider

    Go to Dashboard → Providers, pick a provider (e.g. OpenAI or OpusMax), paste its API key and click Test to check it works. Keys are encrypted (AES-256-GCM) and only the last 4 characters are ever shown. Use Update key to rotate one. Connect several providers to give the router more options.
  3. 3

    Create an API key for each app

    In Dashboard → API Keys, create a key named after the app (e.g. "InstaBizIntel"). Copy the sk-inf-… secret now, because it is shown only once. One key per app gives you per-app logs, filters and budgets.
  4. 4

    Point your app at the gateway

    Set the base URL to https://your-instaroute-host/v1 (OpenAI SDKs) or https://your-instaroute-host (Anthropic SDK / Claude Code) and use your sk-inf-… key in place of the provider key.
  5. 5

    Send requests with model: "auto"

    Use auto or another routing goal. Then open Logs and click any request to see which model was chosen, why, and how much it saved.
Tip:Try the Playground in the dashboard first. It sends a real request through the gateway and shows the routing decision, cost and savings without writing any code.

Connect your app

Any OpenAI- or Anthropic-compatible SDK, framework or tool works. Change the base URL and the key; nothing else in your code needs to change.

Environment variables (most apps)

env
# Most apps that already use OpenAI only need two settings changed:
OPENAI_BASE_URL=https://your-instaroute-host/v1
OPENAI_API_KEY=sk-inf-...        # InstaRoute key, NOT your OpenAI key
OPENAI_MODEL=auto                # or a goal / model id

Python: OpenAI SDK

python
from openai import OpenAI

client = OpenAI(
    base_url="https://your-instaroute-host/v1",
    api_key="sk-inf-...",          # your InstaRoute key (one per app)
)

resp = client.chat.completions.create(
    model="auto",                  # or a goal: "cheapest", "coding", ... or a model id
    messages=[{"role": "user", "content": "Summarise this in one line: ..."}],
)
print(resp.choices[0].message.content)

# Routing metadata (provider, model, cost, savings) is on the raw response:
print(resp.model_extra.get("instaroute"))

Node.js / TypeScript: OpenAI SDK (streaming)

typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://your-instaroute-host/v1",
  apiKey: process.env.INSTAROUTE_API_KEY, // sk-inf-...
});

const resp = await client.chat.completions.create({
  model: "auto",
  messages: [{ role: "user", content: "Hello!" }],
  stream: true,                         // streaming (SSE) works as usual
});
for await (const chunk of resp) process.stdout.write(chunk.choices[0]?.delta?.content ?? "");

cURL

bash
curl https://your-instaroute-host/v1/chat/completions \
  -H "Authorization: Bearer sk-inf-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "messages": [{"role": "user", "content": "Explain recursion simply."}],
    "instaroute": { "deny": ["gpt-5.5-pro"] }
  }'

Anthropic SDK

The /v1/messages endpoint speaks Anthropic's format, so Claude-format requests can be served by any connected provider. Keys work in either the x-api-key or the Authorization: Bearer header. For Claude Code, see Use with Claude Code.

python
# Anthropic SDK (Python)
from anthropic import Anthropic
client = Anthropic(base_url="https://your-instaroute-host", api_key="sk-inf-...")
msg = client.messages.create(
    model="auto",                 # or "coding", or "claude-sonnet-5-5"
    max_tokens=1024,
    messages=[{"role": "user", "content": "Refactor this function ..."}],
)
print(msg.content[0].text)

OpenAI Responses API

python
resp = client.responses.create(       # OpenAI SDK client from above
    model="auto",
    instructions="You are a concise assistant.",
    input="Give me three taglines for a coffee shop.",
)
print(resp.output_text)

Embeddings

python
emb = client.embeddings.create(
    model="text-embedding-3-small",       # pin a model for anything you store
    input=["first document", "second document"],
)
print(len(emb.data[0].embedding))
Note:Frameworks such as LangChain, LlamaIndex, Vercel AI SDK and LiteLLM work through their OpenAI-compatible client. Set the base URL to https://your-instaroute-host/v1 and use your sk-inf-… key.

Use with Claude Code

Claude Code, Anthropic’s coding agent, can send all of its requests through InstaRoute. Every request is then logged with its cost, counts against your key’s budget, and can be routed to a cheaper model. Claude Code talks to POST /v1/messages (streaming and tool use) and /v1/messages/count_tokens, both of which InstaRoute serves.

Before you start

  • Connect Anthropic or OpusMax under Providers so Claude models are available. Other providers are used only if you turn on auto (below).
  • Create a gateway key under API Keys, named for example “Claude Code”, so its usage shows separately in Logs. Give it a monthly budget if you want a hard spending limit.

1. Point Claude Code at InstaRoute

Use the base URL without /v1. ANTHROPIC_AUTH_TOKEN sends your key as Authorization: Bearer and takes priority over a saved claude.ai login immediately; you do not need to log out. (ANTHROPIC_API_KEY also works, but Claude Code asks you to approve it once.)

bash
# 1. Point Claude Code at InstaRoute (add to ~/.bashrc or ~/.zshrc to keep it)
export ANTHROPIC_BASE_URL=https://your-instaroute-host
export ANTHROPIC_AUTH_TOKEN=sk-inf-...            # your InstaRoute gateway key
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1  # no telemetry or update checks

# 2. Start Claude Code as usual
claude

Or set the same values in a settings file, which applies even when you start Claude Code from an IDE:

json
// ~/.claude/settings.json (all projects) or .claude/settings.local.json (one project)
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://your-instaroute-host",
    "ANTHROPIC_AUTH_TOKEN": "sk-inf-...",
    "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1"
  }
}

2. Choose how models are picked

bash
# Default: Claude Code asks for Claude models by name and InstaRoute serves them
# from your Anthropic or OpusMax connection.

# Let InstaRoute choose the model for every request:
export ANTHROPIC_MODEL=auto

# Or pin what Claude Code's model aliases map to:
export ANTHROPIC_DEFAULT_OPUS_MODEL=claude-opus-5-5
export ANTHROPIC_DEFAULT_SONNET_MODEL=claude-sonnet-5-5
export ANTHROPIC_DEFAULT_HAIKU_MODEL=claude-haiku-4-5

# Keep one repository on one model so the prompt cache stays warm:
export ANTHROPIC_CUSTOM_HEADERS="x-instaroute-session-id: my-repo"

With auto, Claude Code prints a one-line “unrecognized model” notice and carries on; InstaRoute picks the model. Claude Code depends heavily on tool calls, so if you allow non-Claude models, restrict auto to tool-capable ones with a routing rule or the allow list.

3. Check it is connected

  • In Claude Code, run /status. It should show Anthropic base URL set to https://your-instaroute-host and the auth token coming from ANTHROPIC_AUTH_TOKEN.
  • Open Dashboard → Logs and filter by the “Claude Code” app. Each step Claude Code takes appears as a request with model, routing reason, tokens and cost.

4. Test the integration

Run these three checks from any project folder. Each should finish without an error and appear in Logs.

bash
# Test 1 - connection: one prompt, no tools
claude -p "Reply with the word ready"

# Test 2 - tool use: Claude Code must read a file and answer from it
echo "The launch code is BLUE-42." > note.txt
claude -p "What does note.txt say? Read the file."

# Test 3 - routing: let InstaRoute pick the model
ANTHROPIC_MODEL=auto claude -p "Summarise what this folder contains"
TestWhat a pass looks like
1. ConnectionClaude Code prints "ready". Logs shows one request from the Claude Code key.
2. Tool useClaude Code reads note.txt and answers "BLUE-42". Logs shows two requests: the tool call and the final answer.
3. RoutingLogs shows requested "auto" and the model the router chose, with its reason.

Troubleshooting

SymptomFix
401 / invalid API keyThe key must be a gateway key (sk-inf-…), not your Anthropic key. Check /status shows ANTHROPIC_AUTH_TOKEN.
404 or connection refusedANTHROPIC_BASE_URL must not end in /v1, and must be reachable from your machine (https://… for a remote server).
Unknown model errorClaude Code asked for a model your connections don't serve. Connect Anthropic/OpusMax, or set ANTHROPIC_MODEL / ANTHROPIC_DEFAULT_*_MODEL to a catalog id or auto.
Errors about beta featuresSet CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1 so Claude Code stops sending pre-release request fields.
Notice about auto mode classifier billingInformational only. Claude Code keeps working; its auto-mode classifier requests are billed as normal requests.
Same answer for a repeated promptAn identical first prompt can be answered from the exact-match response cache. Any change in prompt or files gives a fresh answer.
Important:InstaRoute translates Claude Code’s requests rather than forwarding them byte for byte, so it can route them to any provider. Extended-thinking blocks are not passed back yet, and anthropic-beta headers are not forwarded to the provider.

Endpoints

All inference endpoints need a gateway key: Authorization: Bearer sk-inf-… or x-api-key: sk-inf-….

Method & pathFormatWhat it does
POST /v1/chat/completionsOpenAIChat completions with streaming (SSE), tools / function calling, JSON mode, vision inputs.
POST /v1/messagesAnthropicAnthropic Messages API: system prompts, text and image content, tool_use / tool_result, tools and tool_choice, stop_sequences, streaming events.
POST /v1/messages/count_tokensAnthropicReturns a local token estimate: { "input_tokens": n }.
POST /v1/responsesOpenAIStateless Responses API: input (string or items), instructions, function tools, tool_choice, text.format (JSON / JSON schema), max_output_tokens, streaming.
POST /v1/embeddingsOpenAIEmbeddings from OpenAI, Gemini, Mistral, Cohere, Together, Fireworks or local Ollama. Pass an embedding model id, or "auto". See Embeddings below.
GET /v1/modelsOpenAILists every routable chat model plus embedding models.
POST /v1/route/explainInstaRouteSame body as chat completions; returns how it would be routed (ranked candidates, why others were ruled out) without calling a model. Free, not logged.
GET /health—Liveness check, no auth: { "status": "ok" }.
Important:/v1/responses is stateless. previous_response_id, conversation, background: true and built-in tools (web search, file search, etc.) return 400 with an explanation, so send the full conversation in input instead. store is ignored.

The model field

The model field accepts three kinds of value:

ValueExampleBehaviour
auto"auto"Complexity-adaptive smart routing. Recommended.
A routing goal"cheapest"The router picks the best model for that intent. See the goals below.
A concrete model id"claude-sonnet-4-6"Explicit routing: that model is always tried first. Canonical ids and raw provider ids (e.g. "claude-sonnet-4-5-20250929") both work. If it fails, failover moves to the next best eligible model.

The full live list (ids, providers, context windows, prices) is returned by GET /v1/models and shown in the dashboard Playground model picker.

Embeddings

Send embeddings to POST /v1/embeddings in the OpenAI format, with the same gateway key. The model field takes an embedding model, not a chat model or routing goal:

ValueExampleBehaviour
A model id"gemini-embedding-001"Always that model. Recommended for anything you store in a vector database. Provider-prefixed ("mistral/mistral-embed") and raw provider ids ("BAAI/bge-base-en-v1.5") also work.
auto"auto"The default model of your first connected provider, in this order: OpenAI, Gemini, Mistral, Cohere, Together, Fireworks, Ollama. It does not change when you connect a cheaper provider later. Omitting model does the same.
cheapest"cheapest"The lowest-priced model among your connected providers. Can change when you connect a provider, so avoid it for stored vectors.
python
emb = client.embeddings.create(
    model="gemini-embedding-001",
    input=["Refund policy for damaged items", "How to update a GST number"],
    dimensions=768,                      # only for models marked "adjustable"
)
print(emb.instaroute["dimensions"], emb.instaroute["cost_usd"])

How an embeddings request is handled

  1. Authenticate. The gateway key identifies your workspace and app, checks the plan quota and the key’s monthly budget, and applies its rate limit.
  2. Resolve the model. A named model is used as is. auto takes the default model of your first connected provider (fixed order), and cheapest takes the lowest price. If the model’s provider is not connected, the request stops with a message listing the models you can use.
  3. Call the provider with your own key (BYOK), at the provider’s embeddings endpoint. The input (one string or a batch) and optional dimensions are passed through unchanged.
  4. Return OpenAI-format vectors in input order, plus an instaroute object with the provider, model, vector size, cost and latency.
  5. Log metadata only: token count, cost, latency and status appear in Logs and in your spend totals. The text you embedded is never stored.

Typical retrieval (RAG) setup

  1. Pick one embedding model and write its id in your config.
  2. Embed your documents in batches (send a list in input) and store the vectors in your vector database.
  3. Embed each user question with the same model and search the database.
  4. Send the question plus the matching passages to /v1/chat/completions with model: "auto"; the chat side is routed for cost as usual.

Available models

The gateway did not report any embedding models. Update the gateway to the latest version.

  • Vectors from different models cannot be compared. If you change model, re-embed your stored documents. For the same reason an embeddings request is never failed over to another model.
  • The model’s provider must be connected under Providers. Ollama models must be pulled on your Ollama server first (for example ollama pull nomic-embed-text).
  • dimensions works only on models marked adjustable; other models return 400 if you send it.
  • Every call is logged with its cost under Logs (goal “embeddings”). The response adds instaroute with the provider, model, dimensions, cost and latency. Gemini does not report token usage on this endpoint, so its tokens are estimated (usage_estimated: true).
  • Try it without code in Dashboard → Playground → Embeddings.

Routing goals

GoalWhat it optimisesGood for
autoAuto (recommended)Best overall value. Measures how hard each request is and adjusts: simple prompts go to cheap, fast models; hard ones go to stronger models.Default for almost every app.
balancedBalancedA strong general-purpose model with good quality at reasonable cost, with fixed weights.Predictable general chat / assistants.
cheapestCheapestThe lowest-cost model that can still do the task acceptably.Bulk classification, tagging, extraction, summaries.
fastestFastestThe lowest-latency model, optimising time-to-first-token.Autocomplete, live UI, voice.
qualityHighest qualityThe most capable model available; cost is secondary.Critical answers, final drafts, complex analysis.
codingCodingOnly coding-tagged models. Best for writing, editing and debugging code and following tool specs.Code assistants, agents, Claude Code.
reasoningReasoningOnly reasoning-tagged models. Multi-step reasoning, maths, planning.Math, logic, planning, research agents.
visionVisionOnly models that accept images; picks the best one for the task.Screenshots, documents, photos.
long_contextLong contextOnly long-context models; prefers the largest context window.Long PDFs, transcripts, large codebases.
local_firstLocal firstPrefers a self-hosted Ollama model (zero cost, fully private); uses cloud only when local can't do the task.Private data, offline / zero-cost workloads.

Goals with a required capability (coding, reasoning, vision, long_context, local_first) only consider models tagged with it. Requests that contain images, tools or very long inputs are automatically limited to models that support them, whatever the goal.

Request options (instaroute object)

Add an optional instaroute object to the request body to control routing for that one call. All fields are optional.

FieldTypeEffect
providerstringRoute only to this provider (e.g. "openai", "opusmax").
goalstringRouting goal, as an alternative to putting it in model.
allowstring[]Allow-list: the router may only choose from these model ids.
denystring[]Deny-list: never use these model ids.
no_cachebooleanBypass the exact and semantic response cache for this request.
no_jevbooleanSkip JEV and use the built-in heuristic scorer.
deferrablebooleanThe caller can wait: use the flex tier (~50% off, slower, may queue) where the model offers it. Falls back to standard if flex is busy. Also triggered by OpenAI's own service_tier: "flex".
session_idstringConversation id for prompt-cache sticky routing. Optional: without it InstaRoute derives one from the system prompt + first user message.
json
{
  "model": "auto",
  "messages": [{ "role": "user", "content": "..." }],
  "instaroute": {
    "provider": "opusmax",               // only route to this provider
    "allow": ["claude-sonnet-4-6", "gpt-5.4-mini"],  // choose only among these
    "deny": ["gpt-5.5-pro"],             // never use these
    "no_cache": true,                    // skip the exact + semantic response cache
    "no_jev": true,                      // heuristic scorer only for this call
    "deferrable": true,                  // OK to wait: use the ~50% cheaper flex tier
    "session_id": "chat-8271"            // keep this conversation on one model (prompt cache)
  }
}

The same options work on /v1/messages and /v1/responses (add the instaroute object to the body), or as headers:

http
# Clients that can set headers but not body fields (e.g. Claude Code):
x-instaroute-deferrable: true
x-instaroute-session-id: chat-8271

# Claude Code example
export ANTHROPIC_CUSTOM_HEADERS="x-instaroute-session-id: my-repo"
Note:With the OpenAI Python SDK, pass these via extra_body={"instaroute": {…}}. In Node, add the field to the request object (cast if TypeScript complains).

Response metadata

Every response, including the final streamed chunk, carries an extra instaroute object. Standard SDK fields are left unchanged.

json
"instaroute": {
  "provider": "openai",
  "model": "gpt-5.4-mini",
  "requested": "auto",
  "router": "heuristic",
  "reason": "simple task (complexity 0.18) — best value among 12 candidates",
  "attempts": 1,
  "cached": false,
  "latency_ms": 812,
  "cost_usd": 0.000214,
  "baseline_usd": 0.004650,
  "savings_usd": 0.004436,
  "service_tier": "standard",
  "savings_breakdown": {
    "routing": 0.004436,
    "prompt_cache": 0,
    "tier": 0,
    "response_cache": 0
  }
}
FieldMeaning
providerProvider that served the request.
modelConcrete model used.
requestedWhat you asked for (goal or model id).
routerHow it was chosen: explicit, rule, jev, heuristic or cache.
confidenceJEV confidence, 0–1 (JEV decisions only).
reasonShort, human-readable reason for the choice.
attemptsUpstream attempts (more than 1 means failover happened).
cachedServed from the response cache (cache_type: exact or semantic; semantic_similarity for semantic hits).
latency_msEnd-to-end latency.
cost_usdActual cost of this request, from the provider's token counts.
baseline_usdWhat the most expensive suitable model would have cost.
savings_usdbaseline_usd − cost_usd (never negative).
savings_breakdownWhere the saving came from: routing, prompt_cache, tier, response_cache (USD; they add up to baseline − cost).
service_tierstandard, flex or priority, as reported by the provider.
prompt_cacheHow the prompt-cache optimiser acted: sticky, breakpoints or both.
cached_input_tokensPrompt tokens read from the provider's prompt cache (also in usage.prompt_tokens_details.cached_tokens).
cache_write_tokensPrompt tokens written to the provider's prompt cache (Claude).

Errors & limits

Errors follow the OpenAI shape { error: { message, type, code } }. On /v1/messages they follow the Anthropic shape { type: "error", error: { type, message } }.

HTTPcodeMeaning / fix
400invalid_request_errorMalformed body or unsupported feature. The message says what to change.
401invalid_api_keyMissing, wrong or disabled sk-inf key.
402budget_exceededThis API key hit its monthly budget. Raise or clear the budget under API Keys.
403permission_errorYour role can't do this (dashboard API).
422no_providerNo connected provider can serve this request (e.g. vision requested but no vision model connected). Connect one, or loosen allow / deny / provider.
429rate_limit_exceededMore than 20 requests/s (burst 40) on one key, or the plan's monthly request quota is used up.
502upstream_errorEvery provider attempt failed (after failover). The message contains the provider's error.

Limits: 20 requests/second per key, with bursts up to 40. Monthly request quota depends on your plan (see Plans).

How routing works

Each request goes through these steps, in order:

  1. 1

    Guardrails (if enabled)

    Optional prompt-injection and secret-leak screening, which can block or flag a request. Off unless the server enables it.
  2. 2

    Key settings

    If the request says "auto" and the API key has a goal override, that goal is used. If the key has used 85% of its monthly budget, goal routing is switched to cheapest.
  3. 3

    Find candidates

    Only enabled, priced models on your connected providers. Models that lack a needed capability (vision, tools, context size) and providers marked down are dropped, and the request's allow / deny / provider options and the key's provider allow-list are applied.
  4. 4

    Measure complexity (auto)

    The prompt is scored 0–1 (simple / moderate / complex) from its length, structure, code, maths and reasoning cues. Simple tasks weight cost; complex tasks weight quality.
  5. 5

    Score & rank

    Each candidate gets a score from cost, latency and quality using the goal's weights. Explicit model ids skip this step.
  6. 6

    Routing rules

    If one of your enabled rules matches, its preferred models move to the front. The decision is logged as router "rule".
  7. 7

    Sticky conversation (prompt cache)

    If this conversation's previous turn ran on a model that is still eligible, and the prompt is large enough for provider caching, it stays there unless moving is cheaper even after losing the cache, or the task got harder.
  8. 8

    JEV (optional)

    If JEV is on for the workspace, it reviews the top 5 and picks one. If it isn't confident enough or is slow, the heuristic pick is used.
  9. 9

    Call & fail over

    The request goes to the chosen model. On 408/409/429/5xx or a network error it moves to the next best model, up to 4 attempts in total.
  10. 10

    Measure & log

    Actual cost (from the provider's token counts), baseline cost, savings, latency and the decision reason are logged. Prompt text is not stored.
router valueMeaning
explicitYou named a concrete model.
ruleOne of your routing rules decided.
jevJEV picked the model.
heuristicThe built-in cost / latency / quality scorer picked it.
cacheServed from the response cache, at $0.

Routing explanations

Send any chat request to POST /v1/route/explain to see how it would be routed, without calling a model or paying for tokens. It goes through the same steps as a real request (your connections, key limits, budget, routing rules and benchmark scores) and returns:

  • chosen and failover: the model that would be tried first and the order after it.
  • candidates: every eligible model ranked by score, with its 0–1 cost, speed and quality components, price, latency and quality (and whether that quality came from your benchmark scores).
  • weights, goal and complexity: how much cost, speed and quality counted for this request.
  • rejected: every other model with the reason it was ruled out (provider not connected, key or request restrictions, missing vision or tools, context too small, provider down).
bash
curl https://your-instaroute-host/v1/route/explain \
  -H "Authorization: Bearer sk-inf-..." -H "Content-Type: application/json" \
  -d '{"model": "auto", "messages": [{"role": "user", "content": "Summarise this contract"}]}'

The same view is in Playground → Explain routing. JEV is not consulted in an explanation (it would cost a call), so in a JEV workspace a live request may pick a different one of the top five.

JEV smart routing

JEV (by TypeSafe AI) is a decision model. It reads the request's task profile and picks the best model from the router's top-5 shortlist. Turn it on per workspace in Dashboard → Settings → JEV routing.

SettingEffect
JEV offOnly the heuristic scorer is used. Fast, free and predictable.
JEV on + your keyJEV chooses among the top-5 candidates using your own TypeSafe AI key (BYOK, stored encrypted).
JEV on, no keyUses the server's JEV key, if the operator has set one. Otherwise falls back to the heuristic.
TestChecks that the key works and shows its latency.

JEV is always optional and never blocks a request. If it times out, errors, or its confidence is too low, the heuristic choice is used and the log shows why. Use no_jev: true to skip it for a single call. Only owners and admins can change this setting.

Routing rules

Rules let you enforce your own policy, e.g. "coding goes to Claude Sonnet" or "requests with images go to GPT-5.4 mini". Manage them in Dashboard → Routing rules (owners and admins).

FieldMeaning
nameA label for the rule.
priorityLower number = checked first. The first matching rule wins.
match.goalMatch requests that use this goal (e.g. coding).
match.requiresVisionMatch requests that contain images.
match.requiresToolsMatch requests that define tools / functions.
match.requiresReasoningMatch reasoning-heavy requests.
match.minContextMatch when the estimated input is at least this many tokens.
preferOrdered list of model ids to use. The first one that is eligible is tried first; the others stay available as failover.
http
POST /api/routing/rules          (dashboard → Routing rules does this for you)
{
  "name": "Coding goes to Claude Sonnet",
  "priority": 10,
  "match":  { "goal": "coding" },
  "prefer": ["claude-sonnet-4-6", "gpt-5.4-mini"]
}
Note:All conditions in a rule must match. Rules apply to goal routing (including auto); a request naming a concrete model always gets that model. Changes take effect within about 15 seconds.

Savings Autopilot & caching

Savings Autopilot is a set of cost levers InstaRoute runs on your traffic. Switch them in Dashboard → Settings → Savings Autopilot (owners and admins). Every saving is recorded per request and summed on the Overview under Savings by lever.

LeverDefaultWhat it does
Prompt-cache optimiserOnProviders charge ~10% for repeated prompt prefixes, but only on the same model. InstaRoute keeps a continuing conversation on its model for up to 10 minutes (sticky routing), and adds cache breakpoints to long Claude prompts (tools, system prompt, conversation so far). Writing a Claude cache costs 25% extra once; every later turn reads it at 10%.
Flex tier for deferrable requestsOnRequests marked deferrable are ranked on flex prices and sent with OpenAI's service_tier: "flex" (~50% off, slower, may queue). If flex is busy (429) or the model has no flex tier, InstaRoute retries at standard automatically.
Exact response cacheAlwaysIdentical requests (same messages, model / goal, temperature, max tokens, tools, response format, seed and routing options) within 24 hours are answered from cache at $0, streaming or not.
Semantic cacheOff (Developer plan +)Answers a reworded repeat when everything except the last user message is identical (same app, system prompt, history and parameters) and the last message means the same thing (similarity at or above your threshold, default 0.95). Uses your OpenAI key for small embeddings ($0.02 per 1M tokens), added to the request's cost. A share of hits (default 5%) is answered fresh and compared with the cached answer, so you can see the agreement rate.
  • Caches are per workspace, so two customers never share entries.
  • Requests that use tools, images or several choices are not answered from the response cache.
  • Add "instaroute": { "no_cache": true } to always get a fresh answer.
  • The semantic index lives in the gateway process, so it starts empty after a restart. Only vectors and answers are kept, never prompt text.

Providers (BYOK)

Connect providers in Dashboard → Providers. Each connection has a label, the encrypted key, an optional custom base URL, an enabled switch and a live health status (up / degraded / down, re-checked every few minutes). Use Test to check a key before or after saving, and Update key to rotate it without losing settings.

ProvideridNotes
OpenAIopenaiGPT-5.x, GPT-4.1, o-series, embeddings. Premium tiers (*-pro, o1, o1-pro) are off by default.
AnthropicanthropicClaude models direct from Anthropic.
OpusMaxopusmaxAnthropic-compatible gateway serving Claude models at lower cost (api.opusmax.pro).
Google GeminigoogleGemini models via the Generative Language API.
GroqgroqVery low-latency open models.
DeepSeekdeepseekLow-cost chat and reasoning models.
xAI (Grok)xaiGrok models.
Mistral AImistralMistral / Codestral models.
Together AItogetherHosted open models: GPT-OSS, DeepSeek V4, GLM 4.6, Llama 3.3, Qwen 3.5. BGE embeddings.
Fireworks AIfireworksFast hosted open models: GPT-OSS, DeepSeek V4, GLM 4.6, Kimi K2, Qwen 3.7. Qwen3 embeddings.
CoherecohereCommand A, Command R and R7B via Cohere's OpenAI-compatible API. Embed v4 and multilingual embeddings.
Ollama (local)ollamaSelf-hosted models, $0 per token. Needs OLLAMA_ENABLED=1 on the server.
OpenRouteropenrouterOne key for hundreds of models; GPT-5.4 mini, Gemini 3.8 Flash, Kimi K2.6, MiniMax M3 and GLM-5.3 Flash are pre-priced.
PerplexityperplexitySonar models with live web search. Used only when you ask for a Sonar model or the perplexity provider; per-request search fees are not included in cost.
CerebrascerebrasVery fast inference: GPT-OSS 120B, Qwen 3.8 27B.
SambaNovasambanovaFast open models: GPT-OSS 120B, MiniMax M2.7, Llama 3.3 70B.
DeepInfradeepinfraLow-cost hosted open models: GPT-OSS, DeepSeek V4 Flash, GLM 4.6, Kimi K2.6, Qwen 3.5.
Nebius AInebiusHosted open models: GPT-OSS, DeepSeek V4 Flash, Qwen3 235B, GLM-5.3, Kimi K2.6.
Novita AInovitaHosted open models: GPT-OSS, DeepSeek V4 Flash, GLM 4.6, Kimi K2.6, MiniMax M3.
Moonshot AI (Kimi)moonshotKimi K2.5, K2.6 and K2.7 Code direct from Moonshot.
Z.ai (GLM)zaiGLM-5.3, GLM-5.3 Flash and GLM 4.6 direct from Z.ai.
Alibaba Cloud (Qwen)dashscopeQwen Max, Plus and Turbo via DashScope's OpenAI-compatible mode (international endpoint).
MiniMaxminimaxMiniMax M3 and M2.5 with 1M-token context.
AI21 Labsai21Jamba Large and Mini 1.7 with 256K context.
FriendliAIfriendliServerless GLM-5.3, GLM-5.3 Flash and MiniMax M2.5.
ScalewayscalewayEU-hosted open models. Priced when discovery finds them (runs automatically after you connect).
BasetenbasetenModel APIs: GPT-OSS, DeepSeek V4 Flash, GLM 4.7, Kimi K2.6.
Azure OpenAIazureNeeds your resource URL (https://<resource>.openai.azure.com/openai). Name deployments after their model (e.g. gpt-5.4-mini); a missing deployment fails over to the next model.
Cloudflare Workers AIcloudflareNeeds your account URL (https://api.cloudflare.com/client/v4/accounts/<account_id>/ai) and a Workers AI token.
Custom endpointcustomAny OpenAI-compatible API: vLLM, LM Studio, NVIDIA NIM, TGI or a private gateway. One per workspace. See Custom endpoint below.

The number of providers you can connect depends on your plan. Only connected and enabled providers are used. Providers marked degraded are ranked lower and providers marked down are skipped.

Note:OpenAI, Anthropic, OpusMax, Google Gemini, DeepSeek, xAI, Mistral, Groq and Ollama models come pre-priced and ready to route. New models from those providers are priced automatically by weekly discovery. For Together, Fireworks and Cohere, discovered models appear in Admin → Models, where an admin confirms the price and enables them. Unpriced models are never auto-routed.

Providers added in October 2026 (OpenRouter through Cloudflare in the table) come with their most-used models pre-priced. Their other models are priced only when the id matches the price list exactly, because these hosts sell variants of one model at different prices (for example gpt-oss-120b and gpt-oss-120b-Turbo). Anything else waits in Admin → Models for review.

Custom endpoint

Connect any OpenAI-compatible API, such as vLLM, LM Studio, NVIDIA NIM, Hugging Face TGI or your own gateway, as the Custom endpoint provider. Enter its base URL including the version path (for example https://llm.example.com/v1), an API key if it needs one, and the models it serves with what each costs you per 1M tokens (0 for self-hosted). Fetch models from endpoint fills the list from its /models route.

models
# id, input $/1M, output $/1M, context window, capabilities
meta-llama/Llama-3.3-70B-Instruct, 0, 0, 131072, tools
qwen2.5-coder-32b, 0.2, 0.6
  • The models are available only to your workspace, as custom:<id> (or by their own id), and take part in auto and goal routing like any other model.
  • Prices you enter are used for cost, routing and savings. A model priced at 0 counts its full baseline as savings, as local Ollama models do.
  • For security the endpoint must be a public https:// host; private, loopback and cloud-metadata addresses are refused and redirects are not followed. Self-hosted operators can allow private hosts with ALLOW_PRIVATE_ENDPOINTS=1.
  • One custom endpoint per workspace. Use Edit models on its card to change the list.

Model catalog & pricing

  • Every model has a price per 1M input / output tokens, plus cached-input, cache-write and flex-tier prices where the provider offers them, taken from the providers' price lists. These prices drive routing, cost and savings. Long-context surcharges (e.g. above 200K tokens) and time-of-day discounts are not modelled yet.
  • Premium tiers (e.g. OpenAI *-pro, o1, o1-pro) are listed but disabled for auto-routing by default so they never surprise you on cost. You can still call them by explicit id once an admin enables them.
  • Weekly auto-discovery checks each connected provider for new models. New models matching a known price are priced and enabled automatically; others wait for admin review.
  • Platform admins can review, price, enable or disable any model in Dashboard → Admin → Models, and can run discovery on demand.

API keys & budgets

Create one key per app in Dashboard → API Keys. The page's Connect your app card has ready-to-copy settings. Each key has:

OptionWhat it does
NameShown in logs and filters. Use the app's name.
Monthly budget (USD)At 85% of the budget, goal routing switches to cheapest automatically. At 100%, requests get 402 budget_exceeded until the next month or until you raise the budget.
Goal overrideWhen the app sends model "auto", use this goal instead (e.g. make one app always cheapest). Explicit model ids are not affected.
Allowed providersRestrict this key to certain providers (e.g. OpusMax only).
Disable / deleteDisable blocks the key immediately (reversible). Delete removes it.
Important:The secret is shown only once and stored as a one-way hash. If you lose it, create a new key and delete the old one. Never put sk-inf-… keys in browser or mobile code; call InstaRoute from your server.

Logs & filters

Dashboard → Logs lists every request with time, app, requested goal / model, chosen model, router, tokens, cost, savings, latency and status.

Filters

  • Time: last hour, last 24h, last 7 days, last 30 days, last 90 days, or a custom from–to date range.
  • App: filter by API key.
  • Status: success or error.

Export

Export CSV downloads every request that matches the current filters (up to 50,000 rows) with tokens, cost, baseline and savings by lever. Export FOCUS gives the same data in columns aligned with the FinOps Foundation FOCUS spec (BilledCost, ChargePeriodStart,ServiceCategory …, plus x_ extension columns for baseline and savings) for cost tools such as CloudZero, Vantage or Finout. API: GET /api/logs/export?format=csv|focus&range=30d.

Request detail popup

Click any row to see:

  • Scenario: a metadata-only summary of the task (e.g. "Chat request — 3 messages, ~1200 input tokens, 2 tools, difficulty: moderate.").
  • Routing decision: the baseline (top) model, the model actually used, the router, JEV confidence, complexity score / band, and the reason.
  • Cost & savings: cost if the top model had been used, the actual cost, the amount and percentage saved, and where the saving came from (routing, prompt cache, flex tier, response cache).
  • Request details: tokens in / out, cached and cache-write tokens, latency, app, goal, router, service tier, prompt-cache optimiser, response cache (with semantic similarity and spot-check agreement), and the error code if it failed.

How far back logs go depends on your plan's retention period.

How savings are calculated

Savings are measured per request from real numbers:

formula
baseline cost = price of the MOST EXPENSIVE model that could have served this request
                × the actual input / output tokens of the request
actual cost   = price of the model actually used × the provider-reported tokens
                (at the tier it was served, with cache reads/writes priced separately,
                 plus any embedding cost for the semantic cache)
savings       = max(0, baseline cost − actual cost)

savings by lever (they add up to baseline − actual):
  routing         baseline − chosen model at standard price
  prompt_cache    chosen model without vs with the cache reads InstaRoute produced
                  (net of the one-time cache-write cost, so a write turn can be negative)
  tier            standard price − flex price (a priority-tier request counts negative)
  response_cache  the chosen model's cost avoided by a cache hit, minus embedding cost

Example: a request uses 2,000 input and 500 output tokens. The most expensive suitable model costs $5 / $25 per 1M tokens, so the baseline is 2,000×5/1M + 500×25/1M = $0.0225. The router picked a model at $0.40 / $1.60, so the actual cost is $0.0016. Savings = $0.0209 (93%).

Totals (Overview, Billing and the success fee) are net for the period: the sum of baselines minus the sum of actual costs, never below zero. A request that cost more than its baseline, such as a first Claude turn that pays to write the prompt cache, lowers the total instead of counting as zero, so the total always equals “Spend vs. baseline” and the sum of the levers.

Only models that meet the request's capability needs and your restrictions count toward the baseline. Cache discounts are credited only when InstaRoute produced them (sticky routing or the cache breakpoints it added). Caching a provider does on its own is applied to the baseline too, so it is never counted as an InstaRoute saving. Explicitly requested models are logged with their real cost.

Overview & playground

  • Overview: requests, spend, baseline, savings, cache hit rate, average latency and error rate for the selected range, savings by lever (routing, prompt cache, flex tier, response cache, with prompt-cache tokens, flex requests and semantic spot-check agreement), breakdowns by provider, model and router, and monthly quota usage.
  • Playground: send a test prompt with any goal or model and see the response with its provider, model, router, confidence, latency, cost and savings.
  • Settings: workspace name, default routing goal (used when a request asks for auto), JEV routing and Savings Autopilot.

Spend alerts

Get notified before costs surprise you. Configure in Dashboard → Alerts (owners and admins edit; others can view).

SettingDetails
Workspace monthly budgetThe spend you want to be alerted against.
Stop traffic at the budgetOptional hard limit. From 85% of the budget, goal routing switches to cheapest; at 100% every request gets 402 budget_exceeded until next month or until you raise the budget. Spend is checked every few seconds, so a burst can go slightly over.
ThresholdsPercent levels that trigger alerts. Default 50%, 80%, 100%.
Per-key budget alertsAlert when any API key reaches 80% and 100% of its own budget.
Slack webhookA Slack Incoming Webhook URL (hooks.slack.com).
Generic webhookAny HTTPS URL. Receives a JSON payload, so you can wire it to email, Teams, PagerDuty, etc.
Send testSends a test alert to the configured destinations.

Each threshold fires at most once per month. Recent alerts are listed on the page. Webhook URLs are hidden from non-admins.

Benchmark scores

Every model has a built-in quality score (0–100) that routing weighs against cost and latency. In Dashboard → Benchmarks you can import your own scores, for example from your evals, and they replace the built-in score for your workspace only. Included from the Team plan up; owners and admins import.

csv
model,score,source
gpt-5.4-mini,82,support-evals-2026-10
claude-sonnet-5-5,94,support-evals-2026-10
deepinfra-deepseek-v4-flash,85,support-evals-2026-10
  • CSV (model,score[,source], header optional) or JSON ([{"model": "...", "score": 82}]).
  • Models can be catalog ids or provider model ids. Unknown models and scores outside 0–100 are skipped and listed.
  • New scores merge with existing ones; tick Replace to start over. Changes apply within 15 seconds.
  • When a scored model is chosen, the routing reason in Logs says the quality came from your benchmark scores.

Shadow experiments

Try a model on real traffic before switching to it. In Dashboard → Experiments, pick a candidate model and a sample rate (up to 20%). After a request succeeds and its answer has been returned, InstaRoute sends the same request to the candidate and compares the two. Your users always get the original answer. Included from the Business plan up.

MeasuredHow
Cost / requestEach model's cost from provider-reported tokens, averaged over the samples.
LatencyEnd-to-end time of each call.
Answer agreementSimilarity of the two answers' embeddings (needs an OpenAI connection; costs $0.02 per 1M tokens).
ErrorsCandidate calls that failed or timed out (60 s).
Important:Shadow calls use your provider keys, so they cost what those calls cost. That spend is shown on the Experiments page and is not included in Overview spend or savings. Sampling stops at the sample limit you set. Requests that use tools are not sampled, and only metrics are kept, never prompts or answers.

Team & roles

Invite teammates in Dashboard → Team. You get an invite link to share. It works once, expires after 7 days, and must be accepted by the invited email address (new users register through it). How many members you can have depends on your plan.

RoleCan do
OwnerEverything, including managing other owners. The last owner can't be removed or demoted.
AdminManage providers, routing rules, JEV, alerts, invites and member roles (except owners).
MemberCreate and manage API keys, use the playground, view logs and analytics.
ViewerRead-only: view logs, analytics and settings.

Audit log

Dashboard → Audit log records who changed what: providers connected, updated or removed; API keys; routing rules; benchmark imports; experiments; invites, joins, role changes and removals; workspace, Savings Autopilot and alert settings; and log exports. Owners and admins can view it on every plan, filter by area and time, and click a row for details. CSV export is included from the Team plan up.

Secrets are never recorded: provider keys and webhook URLs appear only as “set” or a key's last four characters.

Plans & billing

You always pay AI providers directly with your own keys. InstaRoute charges a flat plan price, and paid plans add a small success fee on measured savings: if you don't save, you don't pay it. Manage your plan in Dashboard → Billing.

See the pricing section on the homepage for current plans.

Monthly savings statements are under Billing for owners and admins: spend, baseline and net savings for any month, split by cause, day, model and app, with the success fee. Open one to print or save it as PDF, or download it as CSV. API: GET /api/statements/2026-10?format=html|csv|json.

Every new workspace starts with a 7-day free trial of the Developer plan: no card, and no success fee on savings during the trial. When it ends the workspace moves to Free unless you choose a plan. Routing then uses only as many providers as the plan allows (the earliest connected ones). Higher plans add benchmark import, audit export, shadow experiments, white-labelling and priority support. The success fee is calculated from the same per-request savings shown in your logs.

Privacy & security

  • No prompt storage: prompts and responses are never written to logs or the database. Logs keep metadata only (model, tokens, cost, latency, a non-content scenario summary).
  • Encrypted provider keys: BYOK keys (and your JEV key) are encrypted at rest with AES-256-GCM and never returned by the API.
  • Hashed gateway keys: sk-inf-… keys are stored as SHA-256 hashes.
  • Tenant isolation: data and cache entries are scoped to each workspace.
  • Webhook safety: alert webhooks must be HTTPS and are checked so they can't target internal addresses.
  • Local option: use the local_first goal with Ollama to keep data entirely on your own hardware.

Self-hosting settings

For operators running their own InstaRoute server (Docker Compose). Set these in infra/.env; never commit secrets to git.

VariablePurpose
PUBLIC_URLPublic URL of the deployment (used for snippets, CORS and links).
DASHBOARD_URLOptional, only if the dashboard is on a different host (invite and alert links).
MASTER_ENCRYPTION_KEY64 hex chars, used to encrypt provider keys. Required.
JWT_SECRETSigns dashboard sessions. Required.
PLATFORM_ADMIN_EMAILSComma-separated emails that get the Admin panel.
BRAND_NAME / BRAND_TAGLINEWhite-label the product name and tagline.
JEV_API_KEYServer-wide JEV fallback key (workspaces can bring their own).
MODEL_DISCOVERY_ENABLEDWeekly model auto-discovery (default on).
MODEL_DISCOVERY_INTERVAL_DAYSDiscovery interval (default 7).
MODEL_DISCOVERY_AUTO_ENABLEAuto-enable discovered models with known pricing (default on).
OLLAMA_ENABLED / OLLAMA_BASE_URLTurn on local Ollama models.
ALLOW_PRIVATE_ENDPOINTSSet to 1 to let provider base URLs and custom endpoints point at private or loopback hosts (e.g. vLLM on the same server). Off by default.
GUARDRAILS_ENABLED / GUARDRAILS_MODEPrompt-injection and secret screening: block or flag.
STRIPE_*Optional Stripe checkout for paid plans.
bash
docker compose -f infra/docker-compose.yml --env-file infra/.env up -d --build

Platform admin

Users listed in PLATFORM_ADMIN_EMAILS see an Admin section with:

  • Platform totals: workspaces, users, requests, spend, savings, plan MRR and success-fee revenue.
  • All workspaces with owner, plan, usage and projected fee. Change plans or suspend a workspace.
  • Models: the full catalog with price, quality, context, capabilities and status. Enable / disable or edit pricing, and run discovery now.

Troubleshooting & FAQ

I get 401 invalid_api_key.

Use the sk-inf-… key from API Keys, not your OpenAI / Anthropic key. Check the key isn't disabled, and that the header is Authorization: Bearer or x-api-key.

I get 422 no_provider.

No connected provider has a model that can do this request. Connect a provider (and click Test), or loosen the key's allowed providers or the request's allow / provider options.

I get 502 upstream_error.

All attempts failed at the provider. The message contains the provider's error; usually an invalid / expired provider key, no credit, or a model your account can't access. Re-test the provider.

Why did auto pick a model I didn't expect?

Open the request in Logs. The detail popup shows the complexity score, baseline model and the reason. For simple prompts, auto favours cheaper models with good quality, which may be newer or cheaper than the one you'd expect. Use a goal, a routing rule or allow to steer it.

My base URL doesn't work.

OpenAI SDKs need https://your-instaroute-host/v1. The Anthropic SDK / Claude Code needs https://your-instaroute-host without /v1.

Can I force a specific model?

Yes: put its id in model, e.g. claude-sonnet-4-6. Routing rules and goals don't override an explicit model.

Do you see my prompts?

No. Prompts pass through to the provider and are never stored. Only metadata is logged.

Does streaming work?

Yes, on chat completions, messages and responses. Cost and savings are calculated from the final usage.

Ready to route?

Create a free workspace, connect a provider and send your first request in minutes.