Private by architecture · Guarded at every call

One gateway. Every model. Your data stays yours.

Private inference on 200+ models, guarded on every call. Zero markup on pay-as-you-go and subscriptions. Or bring your own keys.

No credit card · live in 60 seconds · prompts aren’t logged unless you turn it on

200+
models, one endpoint
0%
token markup, ever
0 B
prompts stored by default
logging is off until you turn it on
40+%
inference cost savings
with cache-aware adaptive routing
99%
routing accuracy
with frontier escalation
<10ms
routing latency with automatic failover
All integrations
from openai import OpenAI
 
client = OpenAI(
base_url="https://api.orcarouter.ai/v1",
api_key=ORCAROUTER_API_KEY,
)
resp = client.chat.completions.create(
model="orcarouter/auto", # we grade + route
messages=[{"role": "user", "content": "..."}],
)

One line. We grade each prompt, route to frontier or OSS, and add $0.

Privacy & Security

Private by architecture. Guarded at every call.

What you ask an AI stays yours. What your agents do is yours to approve.

Your data

What the model and our servers can see.

  • Data Cloaking

    Cloaked.

    Names, emails and card numbers are swapped for stand-ins before they reach the model you call, and put back in the answer.

    One click
  • ZDR

    Nothing kept.

    Prompts aren’t logged unless you turn logging on. Sign a zero-retention agreement to lock it off.

    Default
  • PQC

    Quantum-safe.

    Logs you choose to keep are sealed with hybrid ML-KEM. The crypto is open source: SCUTTLE.

    Default
  • Data Residency

    Your region.

    Declare US, EU, UK, Asia-Pacific or China, and your compliance reports are tagged with that region.

    On request
  • TEE

    Sealed while it runs.

    Attested enclaves, with proof for every answer.

    On request

Your agents

What gets through, and what runs.

  • Guardrails

    Blocked before it’s billed.

    PII, secrets, jailbreaks and prompt injection stopped before the model sees them. Pattern rules are free; an AI-judge rule bills only its own check.

    One click
  • Agent Firewall

    Every tool call, graded.

    Each tool and MCP call is allowed, sent to a person for review, or blocked before it runs.

    Watching by default
  • Compliance Packs

    Mapped to 32 frameworks.

    One policy per regulation. Watch it on real traffic first, then turn enforcement on.

    Paid plans

Also:Human approval queueThreats tagged to OWASP LLM Top 10 and MITRE ATLASThe public archive of agent incidents

The gateway

Route smarter. Spend less.

Route

Every prompt graded, then sent to the model that answers it best.

Browse models

Observe

Logs when you want them, live spend and a per-prompt receipt.

See a receipt

Manage

Version prompts and reuse cached calls without touching code.

Version a prompt
Models

Every model. One price list.

The newest of 200+ models, live, at the provider's own price.

View all 200+ →
ModelRouted toInput /MOutput /MContextPrivate route
openai/gpt-6.1-solTextNEWOpenAI Direct$2.00$10.001MNo private route
anthropic/claude-sonnet-5.5TextNEWAnthropic Direct$2.00$10.001MNo private route
typesafe/jev-1.13TextNEW—$0.042$0.04266KNo private route
openai/gpt-6-lunaTextNEWOpenAI Direct$0.100$0.5001MNo private route
openai/gpt-6-solTextNEWOpenAI Direct$2.00$10.001MNo private route
anthropic/claude-opus-5.5TextNEWAnthropic Direct$4.00$20.001MNo private route
grok/grok-4.7TextNEW—$2.00$6.00500KNo private route
orca/orcaverify-text1.0TextNEW—$2.00$2.00—No private route
+ 194 more models · prices update every 60 seconds
Setup

Live in 60 seconds.

Step 1

Point your SDK at us

Set base_url to api.orcarouter.ai/v1 and swap your API key. No other code changes needed.

→
Step 2

We route, guard and observe

Graded in under 1ms, with failover, caching and full logs built in.

→
Step 3

You ship, on one endpoint

Direct to each provider at their published rate — we add $0 per token.

Pricing

We never mark up the tokens you buy.

Our revenue comes from optional team features.

Hacker

Free
Forever. Zero markup on all tokens.
Route — 200+ models, auto-failover
Observe — basic dashboard
Manage — prompt versioning
10 API keys · 0% token markup
Start free

Enterprise

Custom
SLA commitments and dedicated capacity.
Everything in Team
Unlimited team seats
Priority access to new models
Every hidden & beta feature
Dedicated infrastructure
99.99% uptime SLA
Dedicated support & custom pricing
Trust & Compliance
Don’t take our word for it. Audit reports are available under NDA, and our post-quantum crypto is open source.
Paying for tokens

Three ways to pay for tokens.

Top-ups and subscriptions bill at provider price. We add $0.

Plans above are for team features. They don't change token price.

From the blog

Fresh from the engine room.

What we're building and why — our latest posts.

All posts →
FAQ

Privacy, security and routing, answered.

Do you store or train on my prompts?

We never train on your prompts. Prompt and response bodies aren't logged unless logging is on for your workspace; we keep request metadata (time, model, tokens, cost, IP) for billing and security. A zero-retention agreement locks body logging off.

Where is my data processed?

By the provider that serves the model you pick, in that provider's regions. Your workspace's data-region setting (US, EU, UK, Asia-Pacific or China) tags your compliance reports with that region; it doesn't yet pin where your data is stored or where inference runs.

What does quantum-safe mean here?

Request logs you choose to keep can be sealed with hybrid post-quantum encryption (ML-KEM-768 + X25519), so a copy taken today can't be opened by a future quantum computer. The code is open source: SCUTTLE. Sealing is on by default on our hosted service.

What is an AI router?

An AI router sits between your app and many models and picks the best model for each request. OrcaRouter grades every prompt in under 1 ms and routes it — frontier for hard reasoning, open-source for routine — across 200+ models at zero token markup.

What is an LLM router and how does it work?

An LLM router classifies each prompt and matches it to the model most likely to answer well at the lowest cost. OrcaRouter routes on contextual embeddings with online learning from live traffic — 75.5% accuracy on the public RouterArena leaderboard.

Is OrcaRouter an AI gateway?

Yes. OrcaRouter is a production AI gateway: one OpenAI-compatible endpoint with adaptive routing, load balancing, automatic failover, guardrails, an agent firewall, prompt caching and per-request observability.

What is an AI CDN?

Like a CDN serves content from the best edge, an AI CDN serves inference from the best provider: healthy, fast, cheap capacity, cached repeated prompts, failover on outages. OrcaRouter plays this role across 200+ models.

What is adaptive routing?

Adaptive routing learns the quality/cost trade-off from your own traffic. Point each workspace at cheapest-that-clears-the-bar, highest quality, or balanced — or let orcarouter/auto keep tuning the choice per request.

Do I need both an AI gateway and an LLM router?

They solve different problems — governance vs model choice — but you don't need two systems: OrcaRouter routes per request and enforces budgets, guardrails and observability on the same hop.

Does routing add latency?

Prompt grading takes under 1 ms and total added latency stays under 50 ms — usually won back many times over by faster providers and cached prompt tokens.

Does OrcaRouter mark up token prices?

No. You pay each provider's published rate; OrcaRouter adds $0 per token and monetizes optional Team and Enterprise features.

Can I keep my existing OpenAI SDK code?

Yes — switch base_url to https://api.orcarouter.ai/v1 and your OpenAI, Anthropic or Google SDK code keeps working: Chat Completions, Responses, Embeddings, Images, Audio and streaming.

What happens when a provider goes down?

OrcaRouter retries against healthy fallback capacity for the same or an equivalent model before the response starts, so upstream outages don't surface to your users.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube