Citation and evidence

OrcaRouter

8 min full readUpdated 23 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI Inference

Cite this article

OrcaRouter is a hosted API gateway and routing service for large language models and other generative models. Applications can address models from multiple providers through one service, select a model explicitly, or delegate the selection to a configured router. Its API includes OpenAI-compatible requests and native Anthropic and Gemini interfaces.[1]

API and model access

The OpenAI-compatible base URL is https://api.orcarouter.ai/v1. An application supplies an OrcaRouter API key rather than the upstream provider's key. The service documents translation between the OpenAI request format and other providers, including support for streaming, tool calling, structured outputs and vision on models that support those capabilities.[2]

Model identifiers normally include a provider prefix. The /v1/models endpoint lists models available to the calling account; the catalog includes endpoint information rather than implying that every model can serve every kind of request. Direct model calls and router calls use the same model field.[3]

The native Anthropic interface uses /v1/messages; its SDK configuration uses the bare host https://api.orcarouter.ai because the SDK adds the path. The native Gemini interface uses /v1beta/models/{model}:generateContent and is documented for the modern google-genai SDK.[4][5]

Routing

Named routers are saved in a workspace and invoked as orcarouter/{name}. Their allowed-model patterns restrict the candidate set, while a default model provides a safety net when the pattern finds no available model. Changing the saved policy changes subsequent routing without changing the application's router identifier.[6]

StrategySelection method
CheapestChooses the lowest-priced live candidate.[6]
QualityChooses the candidate with the highest configured quality score.[6]
BalancedChooses the cheapest candidate meeting a quality threshold, or the highest-quality candidate when none meets it.[6]
Adaptive StandardUses a LinUCB contextual bandit across allowed candidates.[6]
Adaptive GatedNarrows candidates according to request difficulty before applying the adaptive policy.[6]
DSLApplies rules written in YAML and Common Expression Language.[6]

Expanded article table

The orcarouter/auto router is created for an account at signup and remains editable. The current documentation describes a gated-adaptive seed for accounts created from mid-2026 onward, but older accounts may retain a cheapest-model configuration. Adaptive routing initially uses Balanced behavior while accumulating enough observations for a model and request category; its behavior therefore depends on traffic and the saved configuration, not just the alias name.[7]

The routing DSL can match request properties, task classification and agent-session state. Rules are checked in order, with a required default effect when none matches. Destinations can be a model, model list, pool or built-in strategy. The documentation describes dry runs, shadow mode and canary rollout, as well as parallel requests followed by a judge selecting one answer or a synthesizer generating a combined answer.[8]

Fallbacks and conversation continuity

Explicit fallback chains use extra_body.models with extra_body.route set to fallback. The documented limit is five entries, tried in order after upstream failures such as rate limiting, server errors or network errors. A fallback must support the requested endpoint. Once streaming bytes have reached the client, a subsequent failure cannot transparently restart the answer on another model; the client can receive a truncated stream.[9]

Session affinity can associate conversation turns with a model and upstream deployment using X-OrcaRouter-Session-Id. This is intended to maintain continuity and reuse provider prompt caching. Pins are conditional: an unavailable deployment or a model removed from the router's allowed set can cause normal routing to resume. DSL routers re-evaluate their rules instead of using the ordinary model pin.[10]

MCP integration

The official Model Context Protocol server, distributed as @orcarouter/mcp, exposes catalog browsing and chat through MCP clients such as Claude Code, Claude Desktop, Cursor and Windsurf. It requires Node.js 18 or later. Catalog tools work without an API key; the chat tool requires ORCAROUTER_API_KEY.[11]

MCP toolPurpose
orcarouter_models_listLists models and supports provider, capability and minimum-context filters.[11]
orcarouter_model_cardRetrieves details about one model.[11]
orcarouter_providers_listLists providers and model counts.[11]
orcarouter_chatSends a chat request through a model or router.[11]

Expanded article table

The MCP chat implementation accepts a single user prompt and an optional system prompt, defaults to orcarouter/auto, and returns one completed result rather than a stream. Its optional fallback list is combined with the primary model and rejected if the effective chain exceeds five entries.[12]

The MCP server repository is MIT-licensed. That license applies to the published server code; it is not a license for the entire hosted OrcaRouter service.[13]

Routing research

A May 29, 2026 preprint by Continuum AI researchers describes a LinUCB-based router combining lexical features with sentence embeddings. Offline evaluation supplies rewards for every candidate model, initializing one ridge regressor per candidate. Optional online updates use feedback only for the selected model.[14]

For the May 20, 2026 RouterArena submission, the authors report an arena score of 72.08, accuracy of 75.54% and cost of USD 1.00 per 1,000 queries, ranking second at that time. Evaluation used 8,400 queries and a fixed ten-model pool. The deployed submission was evaluated with its policy frozen; separate fits using RouterArena reward matrices were diagnostic, not the submission. The authors also report a robustness score of 22.62 and identify sensitivity to paraphrasing as an improvement area. These results describe that evaluation configuration, not a general service accuracy or price guarantee.[14]

RouterArena evaluates router accuracy, cost and other characteristics across task domains and difficulty levels.[15]

Billing and observability

The billing documentation states that ordinary token usage is charged at upstream provider prices without per-token markup. Provider-side tools can incur additional per-call charges. Responses contain token usage; a request-cost header can request the billed cost, and a request identifier supports later reconciliation. Workspace dashboards aggregate spend, rather than providing a per-key breakdown on the Dashboard itself.[16]

Bring Your Own Key, or BYOK, attaches a workspace's provider credentials so the provider invoices that account directly. On October 7, 2026, the dedicated documentation described an introductory 0% platform fee, with the current rate shown in the console. BYOK is not an unconditional zero-fee promise. By default, an unavailable key can fall back to platform capacity at normal platform prices; enabling Always use this key instead fails the request when that provider's key is unavailable. The documentation limits BYOK to chat-style traffic rather than asynchronous image, video and music jobs.[17]

The service documents a workspace metrics endpoint in OpenMetrics format for request counts, tokens, spend, errors and latency. It requires a key owned by a workspace Admin or Owner. These series cover a rolling 30-day window and are gauges, not monotonic counters; latency generally measures time to first token.[18]

Content policies and data handling

Guardrails are workspace content policies that can block, mask or flag matching request and response text. Built-in checks use string or regular-expression matching; advanced checks can call a model or external vendor. The documentation explicitly describes policy-resolution errors as fail-open: a transient resolution failure can result in no guardrail enforcement.[19]

The Agent Firewall is a separate policy system for tool actions observed at the gateway. It documents checks on advertised tools, emitted tool calls, governed MCP dispatches and reported network destinations. Policies can allow, audit, deny, sanitize, hold for approval or cap cost. Shadow mode records decisions without enforcing them. This does not mean that tools or skills installed outside the gateway are inspected at installation time.[20]

According to the data-handling documentation, prompt and response bodies are not stored by default, but request metadata such as model, token counts, status, latency and source IP is retained. Optional request logging, session replay capture and guardrail raw-text capture retain additional content. Asynchronous media jobs also store results until collection. Turning off prompt logging is therefore different from retaining no data at all.[21]

The gateway processes requests in memory and forwards them to upstream model providers under those providers' terms. BYOK changes which provider account serves the request, not the fact that the request passes through OrcaRouter. The workspace compliance-region setting applies to report artifacts: it does not by itself move request logs or constrain the region where inference runs.[22][23]

References

  1. ^OrcaRouter. Introduction. Accessed October 7, 2026.
  2. ^OrcaRouter. OpenAI SDK compatibility. Accessed October 7, 2026.
  3. ^OrcaRouter. Models. Accessed October 7, 2026.
  4. ^OrcaRouter. Anthropic SDK compatibility. Accessed October 7, 2026.
  5. ^OrcaRouter. Google GenAI SDK compatibility. Accessed October 7, 2026.
  6. ^1 ^2 ^3 ^4 ^5 ^6 ^7OrcaRouter. Named Routers. Accessed October 7, 2026.
  7. ^OrcaRouter. Auto Router. Accessed October 7, 2026.
  8. ^OrcaRouter. Routing DSL. Accessed October 7, 2026.
  9. ^OrcaRouter. Model Fallbacks. Accessed October 7, 2026.
  10. ^OrcaRouter. Session Affinity. Accessed October 7, 2026.
  11. ^1 ^2 ^3 ^4 ^5Continuum AI. OrcaRouter MCP server README. Accessed October 7, 2026.
  12. ^Continuum AI. MCP chat implementation. Accessed October 7, 2026.
  13. ^Continuum AI. MCP server license. Accessed October 7, 2026.
  14. ^1 ^2Bao, Zhenghua, et al. OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning. Preprint, May 29, 2026.
  15. ^Lu, Yifan, et al. RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers. Preprint, revised November 27, 2025.
  16. ^OrcaRouter. Billing and usage. Accessed October 7, 2026.
  17. ^OrcaRouter. Bring Your Own Key. Accessed October 7, 2026.
  18. ^OrcaRouter. Observability. Accessed October 7, 2026.
  19. ^OrcaRouter. Guardrails. Accessed October 7, 2026.
  20. ^OrcaRouter. Firewall. Accessed October 7, 2026.
  21. ^OrcaRouter. Data Handling. Accessed October 7, 2026.
  22. ^OrcaRouter. Data Flow and Trust Boundaries. Accessed October 7, 2026.
  23. ^OrcaRouter. Set data residency. Accessed October 7, 2026.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 1,576 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent full-article review against 23 primary and research references, October 7, 2026. Checked subject identity, specifications, availability, limitations and citation support.

Cite this page: AI Wiki. "OrcaRouter." aiwiki.ai, updated 6 Oct 2026, fact-checked 6 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/orcarouter

Suggest edit