# LLM Model Routing

> Source: https://aiwiki.ai/wiki/llm_model_routing
> Updated: 2026-09-27
> Fact-checked: 2026-09-27
> Categories: AI Inference, Large Language Models, MLOps
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "LLM Model Routing." aiwiki.ai, 27 Sept 2026. https://aiwiki.ai/wiki/llm_model_routing
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

LLM model routing is the selection of a [large language model](https://aiwiki.ai/wiki/large_language_model) to handle a request from a pool of candidate models. A router can choose a model before generating an answer, or participate in a system that tries models sequentially. The objective may be lower cost at an acceptable quality level, higher quality within a budget, or compliance with operational constraints. Routing research evaluates the selection policy as well as the models it can call.[19]

A router does not make its chosen model intrinsically more accurate. Its potential benefit comes from differences between models: a less expensive model may answer some requests adequately, and models can have complementary strengths. The difficult part is predicting which choice will work before the correct answer is known.[8]

## What routing selects

The word "routing" refers to several different decisions in AI systems. They can coexist in one application, but their evaluation criteria differ.

| Mechanism | Decision | Distinction |
|---|---|---|
| Model routing | Which complete model should answer this request? | Changes the model whose capabilities and behavior the application receives.[1] |
| Provider routing | Which provider or endpoint should serve a selected model? | May keep the model identity fixed while changing hosting, price, availability, or supported parameters.[5] |
| Load balancing | Which available deployment should receive traffic? | Distributes operational load; it need not predict answer quality.[6] |
| Model fallback | Which alternative should be tried when an earlier attempt fails? | Usually responds to an error or availability problem, rather than predicting the best model in advance.[18] |
| Quality-based cascade | Should an answer be accepted or should another model be called? | Uses information from generated answers and can incur several generation costs.[2] |
| [Mixture of experts](https://aiwiki.ai/wiki/mixture_of_experts) routing | Which expert components inside a neural network should process an input? | Selects internal computation rather than independently deployed language models.[4] |

For example, an application can first select a language model, then select a hosting provider, and finally apply a fallback rule if that provider fails. Reporting only that the application has a "router" leaves these separate choices unspecified.[5][18]

## Routing policies

A routing policy maps information about a request to a model choice. The information may include the prompt, conversation context, an application-defined task category, or estimates of candidate performance. The policy and its candidate pool are separate: changing either can change the system's results.[19]

| Approach | Basis of selection | Example or limitation |
|---|---|---|
| Explicit task policy | Map a recognized domain or action to a permitted model | Arch-Router studies matching requests to user-defined domain-action preferences.[9] |
| Retrieval-based routing | Estimate performance from similar previously evaluated queries | ContextualRouter uses retrieved historical performance; a nearest-neighbor baseline was competitive in its experiments.[10] |
| Learned performance or preference predictor | Predict model suitability from the input | RouteLLM learns relative preferences; RouterDC learns representations of queries and candidate models.[1][8] |
| Managed quality-cost routing | Use a service's predictor and configured quality criterion | Amazon Bedrock intelligent prompt routing predicts response quality before forwarding the request.[16] |
| Answer-dependent cascade | Score an initial answer, then stop or continue | FrugalGPT learns a sequence of models and acceptance thresholds.[2] |

These approaches answer different questions. A preference-matching router can correctly follow a policy without demonstrating that its chosen model is the most accurate. Likewise, a predictor trained on conversational preferences is not automatically a predictor of code-test success. The label used for training defines what the router is learning.[9][12]

## Training, thresholds, and calibration

Training data normally connects queries with evidence about candidate models, such as task scores or pairwise preferences. RouterDC uses a query encoder, learned model embeddings, and two contrastive objectives. Its purpose is to learn useful query-model relationships, including cases in which more than one candidate performs well.[8]

RouteLLM, introduced as a 2024 preprint and published at ICLR 2025, studies selection between stronger and weaker model classes using preference data. Its methods include similarity-weighted ranking, matrix factorization, a BERT classifier, and a causal language-model classifier. A predicted probability that the stronger model wins is compared with a threshold: raising the threshold sends fewer requests to that model.[1]

Threshold selection is an operational choice, not a universal constant. The RouteLLM implementation includes a calibration utility for targeting a specified fraction of stronger-model calls and recommends calibration on representative queries. A threshold producing a particular call share on one dataset need not reproduce that share on another workload.[7]

This use of "calibration" differs from probability calibration. Probability calibration asks whether events assigned a particular probability occur at approximately that frequency. Selecting a threshold to meet a call budget does not, by itself, establish that a router's scores are calibrated probabilities. Guo and colleagues' work on neural-network calibration also distinguishes confidence estimates from classification accuracy.[11]

The same distinction applies to the predicted event. "The expensive model is preferred" is not the same event as "the inexpensive model is wrong." Both answers might be acceptable, or both might fail. A threshold optimized for relative preference therefore needs an application-level quality check before it can be treated as an acceptance rule.[1][12]

Adding candidates creates another issue: a router trained with a fixed set of model identities may need new performance data or retraining. Retrieval-based methods can instead incorporate historical scores for additional candidates. ContextualRouter, published in the EACL 2026 Student Research Workshop, studies this setting without retraining the router for each change in the model pool. This does not remove the need for evidence about the new model.[10]

## Cost and latency

For token-priced services, request cost depends on input length, generated output length, and the applicable rates, not merely on the number of calls. Cascade cost includes the models invoked before the accepted answer. FrugalGPT explicitly formulates selection under an average-cost budget and accounts for the calls along the cascade.[2]

Operational overhead can include the router itself, embedding generation, retries, and additional deployments. RouteLLM's documented matrix-factorization and similarity-ranking configurations use an embedding API even when the answering models are hosted elsewhere. A locally hosted answer model therefore does not imply an entirely local routing pipeline.[7]

Latency requires its own measurement. A sequential cascade waits for an answer before deciding whether to escalate. Infrastructure routing also encounters queues, timeouts, and rate limits. LiteLLM, for example, exposes latency-based and rate-limit-aware strategies as operational alternatives.[6]

The 2026 LLMRouterBench study illustrates a further distinction: its latency analysis estimates response time from token usage and provider statistics. The authors describe those estimates as indicative and configuration-dependent, not direct measurements of every routed request. Equal estimated cost and accuracy do not imply equal response time.[13]

### Hypothetical cost calculation

The following is an arithmetic illustration, not a benchmark or current provider pricing. Assume 1,000 requests, a small-model cost of $0.002 per request, a large-model cost of $0.020, and routing overhead of $0.0002 per request. Assume these fixed costs already account for token usage.

| Policy | Calculation | Total |
|---|---|---|
| Always use the large model | 1,000 x $0.020 | $20.00 |
| Select the small model for 70% of requests | 700 x $0.002 + 300 x $0.020 + 1,000 x $0.0002 | $7.60 |
| Call the small model first, then escalate 30% | 1,000 x $0.002 + 300 x $0.020 + 1,000 x $0.0002 | $8.20 |

For the last row, the $0.0002 charge is assumed to include answer scoring. The respective reductions against the $20 baseline are 62% and 59%. Neither calculation establishes equal answer quality, acceptable latency, or actual savings for a deployed service. Those require measurement; the example isolates why paying for an initial generation changes cascade economics.

## Evaluating a router

[LLM evaluation](https://aiwiki.ai/wiki/llm_evaluation) for a routed system must test the final answers, not only whether the classifier assigned plausible task labels. RouterBench supports comparison using stored model outcomes, while LLMRouterBench standardizes model pools, datasets, and evaluation adapters. Their controlled comparisons help separate a routing algorithm's contribution from the capabilities of the models available to it.[3][13]

Useful comparisons include the best fixed model on the workload and random selection under comparable resource constraints. An oracle that selects the best observed answer with hindsight estimates available headroom; it is not a deployable router because it already knows the outcomes. RouterBench's oracle selects the cheapest option when several models share the best score.[3]

| Question | Measurement or check |
|---|---|
| Does routing improve the chosen objective? | Compare answer quality and cost across thresholds, including fixed-model baselines.[3] |
| Are probabilities meaningful? | Compare predicted confidence with observed event frequencies on held-out data.[11] |
| Does the scorer measure the intended quality? | Use task-appropriate grading and inspect judge bias, rather than treating all preference scores as correctness.[12] |
| Does performance survive a different task mix? | Report results by domain and difficulty; test changed or unfamiliar inputs.[15] |
| Does the service remain responsive? | Measure latency and failure behavior, including retries and queueing.[6] |
| Are restrictions actually enforced? | Check eligible providers, required parameters, and data-handling settings throughout routing and fallback.[5][18] |

An [LLM-as-a-judge](https://aiwiki.ai/wiki/llm_as_a_judge) can supply scalable preference labels, but those labels have limits. The MT-Bench study documented position, verbosity, and self-enhancement biases, along with weaknesses in judging some reasoning tasks. A routing policy rewarded by that judge can learn the judge's preferences instead of improving the application's intended outcome.[12]

### Interpreting research results

FrugalGPT reported reductions of up to 98% relative to the best individual model in its experiments while matching performance. Its evaluated tasks and historical API costs delimit that result; it is not a promised reduction for arbitrary applications or current prices.[2]

Later work provides reasons to compare simple alternatives. LLMRouterBench, first released in January 2026 and published in Findings of ACL 2026, evaluated 10 routing baselines using more than 400,000 instances spanning 21 datasets and 33 models. Several methods failed to reliably beat its best-single-model baseline. Its coverage did not include all routers, very long-context workloads, or multimodal benchmarks.[13]

The ContextualRouter study similarly found that a simple nearest-neighbor averaging baseline performed comparably to or better than more elaborate alternatives in its experiments. These results do not show that routing never helps. They show that the candidate pool, workload, and baseline matter, and that architectural complexity is not sufficient evidence of improvement.[10]

## Distribution shift and security

[Distribution shift](https://aiwiki.ai/wiki/distribution_shift) can alter the relationship between a prompt and the model that handles it well. The study *How Robust Are Router-LLMs?* examined routing across task categories and difficulty levels. It found cases where tested routers selected a stronger model for straightforward coding or mathematics tasks, and cases where jailbreak prompts were sent to weaker models. These findings concern the evaluated routers and configurations, not every routing system.[15]

Routing is also an attack surface separate from answer generation. *Rerouting LLM Routers* demonstrated input additions that caused evaluated open-source and commercial routers to select more expensive models. An attacker could thereby affect resource allocation even when answer quality was not degraded. The authors call the broader requirement "LLM control plane integrity": preserving orchestration behavior under adversarial input.[14]

Cost attacks differ from attempts to elicit prohibited content, although both can involve manipulated prompts. Testing only the final answer can miss an unwanted change in model selection. A routing assessment therefore needs to examine its decisions and costs as well as generated text.[14]

Data handling belongs in eligibility rules. Provider selection can change where a request is processed, and the router or embedding service may itself receive prompt data. OpenRouter documents separate controls for provider data collection, zero-data-retention endpoints, and support for requested parameters. These controls illustrate why "choose the cheapest model" is an incomplete policy for requests with handling restrictions.[5][7]

## Implementations and operational choices

[OpenRouter](https://aiwiki.ai/wiki/openrouter) separates automatic model selection from provider selection. Its Auto Router documentation, reviewed in September 2026, describes task classification followed by selection using recent aggregate spending patterns and configured restrictions. That market-based signal is different from a directly measured probability of factual correctness.[17]

Amazon Bedrock intelligent prompt routing uses predicted response quality and configurable criteria within a selected model family. Its responses identify the model used, allowing applications to observe the actual choice rather than treating the router's identifier as the answering model.[16]

Fallback configuration needs comparable visibility. OpenRouter's model-fallback documentation describes trying alternatives when a selected model encounters an error and returning the model that served the response. A successful fallback is evidence of request completion, not proof that the replacement matched the original model's quality on the task.[18]

A routing system's operating point should therefore be stated as a tested policy with a defined model pool, workload, quality measure, and cost accounting. A provider's supported options explain how to configure a system; workload-specific evaluation establishes whether that configuration achieves its intended result.[16][19]

## References

1. Ong, Isaac, et al. [RouteLLM: Learning to Route LLMs with Preference Data](https://arxiv.org/abs/2406.18665). arXiv:2406.18665, 2024; revised February 23, 2025. Published at [ICLR 2025](https://proceedings.iclr.cc/paper_files/paper/2025/hash/5503a7c69d48a2f86fc00b3dc09de686-Abstract-Conference.html) under the title *RouteLLM: Learning to Route LLMs from Preference Data*.
2. Chen, Lingjiao, Matei Zaharia, and James Zou. [FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance](https://arxiv.org/abs/2305.05176). arXiv:2305.05176, May 9, 2023.
3. Hu, Qitian Jason, et al. [RouterBench: A Benchmark for Multi-LLM Routing System](https://arxiv.org/abs/2403.12031). arXiv:2403.12031, March 2024.
4. Fedus, William, Barret Zoph, and Noam Shazeer. [Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity](https://arxiv.org/abs/2101.03961). 2021; revised 2022, Journal of Machine Learning Research.
5. OpenRouter. [Provider Routing](https://openrouter.ai/docs/guides/routing/provider-selection). Documentation, accessed September 27, 2026.
6. LiteLLM. [Router: Load Balancing](https://docs.litellm.ai/docs/routing). Documentation, accessed September 27, 2026.
7. LMSYS. [RouteLLM repository and usage documentation](https://github.com/lm-sys/RouteLLM). Accessed September 27, 2026.
8. Chen, Shuhao, et al. [RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models](https://arxiv.org/abs/2409.19886). NeurIPS 2024.
9. Tran, Co, Salman Paracha, Adil Hafeez, and Shuguang Chen. [Arch-Router: Aligning LLM Routing with Human Preferences](https://arxiv.org/abs/2506.16655). arXiv:2506.16655, June 19, 2025.
10. Varangot-Reille, Clovis, Christophe Bouvard, and Antoine Gourru. [Generalising LLM Routing using Past Performance Retrieval: A Few-Shot Router is Sufficient](https://aclanthology.org/2026.eacl-srw.22/). EACL 2026 Student Research Workshop, pp. 304-319, March 2026.
11. Guo, Chuan, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. [On Calibration of Modern Neural Networks](https://proceedings.mlr.press/v70/guo17a.html). ICML 2017, PMLR 70, pp. 1321-1330.
12. Zheng, Lianmin, et al. [Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena](https://arxiv.org/abs/2306.05685). NeurIPS 2023 Datasets and Benchmarks Track.
13. Li, Hao, et al. [LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing](https://aclanthology.org/2026.findings-acl.1881/). Findings of ACL 2026. [Preprint](https://arxiv.org/abs/2601.07206), January 12, 2026.
14. Shafran, Avital, Roei Schuster, Thomas Ristenpart, and Vitaly Shmatikov. [Rerouting LLM Routers](https://arxiv.org/abs/2501.01818). arXiv:2501.01818, January 3, 2025.
15. Kassem, Aly M., Bernhard Schölkopf, and Zhijing Jin. [How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities](https://aclanthology.org/2026.eacl-long.351/). EACL 2026, pp. 7496-7507. [Preprint](https://arxiv.org/abs/2504.07113), 2025.
16. Amazon Web Services. [Understanding intelligent prompt routing in Amazon Bedrock](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-routing.html). Documentation, accessed September 27, 2026.
17. OpenRouter. [Auto Router](https://openrouter.ai/docs/guides/routing/routers/auto-router). Documentation, accessed September 27, 2026.
18. OpenRouter. [Model Fallbacks](https://openrouter.ai/docs/guides/routing/model-fallbacks). Documentation, accessed September 27, 2026.
19. Moslem, Yasmin, and John D. Kelleher. [Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey](https://arxiv.org/abs/2603.04445). Transactions on Machine Learning Research, 2026; arXiv:2603.04445, revised August 30, 2026.

