InstaRoute

Precision-calibrated inference routing — adaptive model arbitration across every provider

The instant AI model router — adaptive model arbitration in milliseconds.

One OpenAI-compatible endpoint across every provider. InstaRoute routes each request to the best model for cost, latency or quality — you bring your own keys, we never mark up a token.

Everything a gateway should be

⚡

OpenAI-compatible

Drop-in /v1/chat/completions. Swap the base URL and key — every existing OpenAI SDK just works.

🔑

BYOK, zero markup

Bring your own provider keys. You pay providers directly at list price — we never mark up tokens.

🧭

Adaptive model arbitration

Complexity-scored, multi-objective model selection in milliseconds — balances cost, latency and capability headroom per request without static model assignments.

🛡️

Automatic failover

When a provider degrades or errors, requests re-route to the next best eligible model, transparently.

💾

Savings Autopilot

Prompt-cache-aware routing, a flex tier for work that can wait, and exact + semantic response caching, with every saving itemised per request.

📊

Observability

Per-request logs, cost-vs-baseline savings, provider/model/router breakdowns. Prompt text is never stored.

Swap one line. Keep your SDK.

InstaRoute speaks the OpenAI wire format. Point your client at our base URL, use ask-inf-…key, and ask for a routing goal like auto, cheapest or quality instead of a fixed model.

python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.instaroute.ai/v1",
    api_key="sk-inf-…",   # your InstaRoute gateway key
)

# Specify a routing objective — the engine arbitrates the optimal model.
resp = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Summarise the quarterly report."}],
)
print(resp.choices[0].message.content)

Start routing in minutes

Create a workspace, connect a provider key, and send your first request through InstaRoute.