Citation and evidence

LLM Model Routing

13 min full readUpdated 19 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI InferenceLarge Language ModelsMLOps

Cite this article

LLM model routing is the selection of a large language model to handle a request from a pool of candidate models. A router can choose a model before generating an answer, or participate in a system that tries models sequentially. The objective may be lower cost at an acceptable quality level, higher quality within a budget, or compliance with operational constraints. Routing research evaluates the selection policy as well as the models it can call.[19]

A router does not make its chosen model intrinsically more accurate. Its potential benefit comes from differences between models: a less expensive model may answer some requests adequately, and models can have complementary strengths. The difficult part is predicting which choice will work before the correct answer is known.[8]

What routing selects

The word "routing" refers to several different decisions in AI systems. They can coexist in one application, but their evaluation criteria differ.

MechanismDecisionDistinction
Model routingWhich complete model should answer this request?Changes the model whose capabilities and behavior the application receives.[1]
Provider routingWhich provider or endpoint should serve a selected model?May keep the model identity fixed while changing hosting, price, availability, or supported parameters.[5]
Load balancingWhich available deployment should receive traffic?Distributes operational load; it need not predict answer quality.[6]
Model fallbackWhich alternative should be tried when an earlier attempt fails?Usually responds to an error or availability problem, rather than predicting the best model in advance.[18]
Quality-based cascadeShould an answer be accepted or should another model be called?Uses information from generated answers and can incur several generation costs.[2]
Mixture of experts routingWhich expert components inside a neural network should process an input?Selects internal computation rather than independently deployed language models.[4]

Expanded article table

For example, an application can first select a language model, then select a hosting provider, and finally apply a fallback rule if that provider fails. Reporting only that the application has a "router" leaves these separate choices unspecified.[5][18]

Routing policies

A routing policy maps information about a request to a model choice. The information may include the prompt, conversation context, an application-defined task category, or estimates of candidate performance. The policy and its candidate pool are separate: changing either can change the system's results.[19]

ApproachBasis of selectionExample or limitation
Explicit task policyMap a recognized domain or action to a permitted modelArch-Router studies matching requests to user-defined domain-action preferences.[9]
Retrieval-based routingEstimate performance from similar previously evaluated queriesContextualRouter uses retrieved historical performance; a nearest-neighbor baseline was competitive in its experiments.[10]
Learned performance or preference predictorPredict model suitability from the inputRouteLLM learns relative preferences; RouterDC learns representations of queries and candidate models.[1][8]
Managed quality-cost routingUse a service's predictor and configured quality criterionAmazon Bedrock intelligent prompt routing predicts response quality before forwarding the request.[16]
Answer-dependent cascadeScore an initial answer, then stop or continueFrugalGPT learns a sequence of models and acceptance thresholds.[2]

Expanded article table

These approaches answer different questions. A preference-matching router can correctly follow a policy without demonstrating that its chosen model is the most accurate. Likewise, a predictor trained on conversational preferences is not automatically a predictor of code-test success. The label used for training defines what the router is learning.[9][12]

Training, thresholds, and calibration

Training data normally connects queries with evidence about candidate models, such as task scores or pairwise preferences. RouterDC uses a query encoder, learned model embeddings, and two contrastive objectives. Its purpose is to learn useful query-model relationships, including cases in which more than one candidate performs well.[8]

RouteLLM, introduced as a 2024 preprint and published at ICLR 2025, studies selection between stronger and weaker model classes using preference data. Its methods include similarity-weighted ranking, matrix factorization, a BERT classifier, and a causal language-model classifier. A predicted probability that the stronger model wins is compared with a threshold: raising the threshold sends fewer requests to that model.[1]

Threshold selection is an operational choice, not a universal constant. The RouteLLM implementation includes a calibration utility for targeting a specified fraction of stronger-model calls and recommends calibration on representative queries. A threshold producing a particular call share on one dataset need not reproduce that share on another workload.[7]

This use of "calibration" differs from probability calibration. Probability calibration asks whether events assigned a particular probability occur at approximately that frequency. Selecting a threshold to meet a call budget does not, by itself, establish that a router's scores are calibrated probabilities. Guo and colleagues' work on neural-network calibration also distinguishes confidence estimates from classification accuracy.[11]

The same distinction applies to the predicted event. "The expensive model is preferred" is not the same event as "the inexpensive model is wrong." Both answers might be acceptable, or both might fail. A threshold optimized for relative preference therefore needs an application-level quality check before it can be treated as an acceptance rule.[1][12]

Adding candidates creates another issue: a router trained with a fixed set of model identities may need new performance data or retraining. Retrieval-based methods can instead incorporate historical scores for additional candidates. ContextualRouter, published in the EACL 2026 Student Research Workshop, studies this setting without retraining the router for each change in the model pool. This does not remove the need for evidence about the new model.[10]

Cost and latency

For token-priced services, request cost depends on input length, generated output length, and the applicable rates, not merely on the number of calls. Cascade cost includes the models invoked before the accepted answer. FrugalGPT explicitly formulates selection under an average-cost budget and accounts for the calls along the cascade.[2]

Operational overhead can include the router itself, embedding generation, retries, and additional deployments. RouteLLM's documented matrix-factorization and similarity-ranking configurations use an embedding API even when the answering models are hosted elsewhere. A locally hosted answer model therefore does not imply an entirely local routing pipeline.[7]

Latency requires its own measurement. A sequential cascade waits for an answer before deciding whether to escalate. Infrastructure routing also encounters queues, timeouts, and rate limits. LiteLLM, for example, exposes latency-based and rate-limit-aware strategies as operational alternatives.[6]

The 2026 LLMRouterBench study illustrates a further distinction: its latency analysis estimates response time from token usage and provider statistics. The authors describe those estimates as indicative and configuration-dependent, not direct measurements of every routed request. Equal estimated cost and accuracy do not imply equal response time.[13]

Hypothetical cost calculation

The following is an arithmetic illustration, not a benchmark or current provider pricing. Assume 1,000 requests, a small-model cost of $0.002 per request, a large-model cost of $0.020, and routing overhead of $0.0002 per request. Assume these fixed costs already account for token usage.

PolicyCalculationTotal
Always use the large model1,000 x $0.020$20.00
Select the small model for 70% of requests700 x $0.002 + 300 x $0.020 + 1,000 x $0.0002$7.60
Call the small model first, then escalate 30%1,000 x $0.002 + 300 x $0.020 + 1,000 x $0.0002$8.20

Expanded article table

For the last row, the $0.0002 charge is assumed to include answer scoring. The respective reductions against the $20 baseline are 62% and 59%. Neither calculation establishes equal answer quality, acceptable latency, or actual savings for a deployed service. Those require measurement; the example isolates why paying for an initial generation changes cascade economics.

Evaluating a router

LLM evaluation for a routed system must test the final answers, not only whether the classifier assigned plausible task labels. RouterBench supports comparison using stored model outcomes, while LLMRouterBench standardizes model pools, datasets, and evaluation adapters. Their controlled comparisons help separate a routing algorithm's contribution from the capabilities of the models available to it.[3][13]

Useful comparisons include the best fixed model on the workload and random selection under comparable resource constraints. An oracle that selects the best observed answer with hindsight estimates available headroom; it is not a deployable router because it already knows the outcomes. RouterBench's oracle selects the cheapest option when several models share the best score.[3]

QuestionMeasurement or check
Does routing improve the chosen objective?Compare answer quality and cost across thresholds, including fixed-model baselines.[3]
Are probabilities meaningful?Compare predicted confidence with observed event frequencies on held-out data.[11]
Does the scorer measure the intended quality?Use task-appropriate grading and inspect judge bias, rather than treating all preference scores as correctness.[12]
Does performance survive a different task mix?Report results by domain and difficulty; test changed or unfamiliar inputs.[15]
Does the service remain responsive?Measure latency and failure behavior, including retries and queueing.[6]
Are restrictions actually enforced?Check eligible providers, required parameters, and data-handling settings throughout routing and fallback.[5][18]

Expanded article table

An LLM-as-a-judge can supply scalable preference labels, but those labels have limits. The MT-Bench study documented position, verbosity, and self-enhancement biases, along with weaknesses in judging some reasoning tasks. A routing policy rewarded by that judge can learn the judge's preferences instead of improving the application's intended outcome.[12]

Interpreting research results

FrugalGPT reported reductions of up to 98% relative to the best individual model in its experiments while matching performance. Its evaluated tasks and historical API costs delimit that result; it is not a promised reduction for arbitrary applications or current prices.[2]

Later work provides reasons to compare simple alternatives. LLMRouterBench, first released in January 2026 and published in Findings of ACL 2026, evaluated 10 routing baselines using more than 400,000 instances spanning 21 datasets and 33 models. Several methods failed to reliably beat its best-single-model baseline. Its coverage did not include all routers, very long-context workloads, or multimodal benchmarks.[13]

The ContextualRouter study similarly found that a simple nearest-neighbor averaging baseline performed comparably to or better than more elaborate alternatives in its experiments. These results do not show that routing never helps. They show that the candidate pool, workload, and baseline matter, and that architectural complexity is not sufficient evidence of improvement.[10]

Distribution shift and security

Distribution shift can alter the relationship between a prompt and the model that handles it well. The study How Robust Are Router-LLMs? examined routing across task categories and difficulty levels. It found cases where tested routers selected a stronger model for straightforward coding or mathematics tasks, and cases where jailbreak prompts were sent to weaker models. These findings concern the evaluated routers and configurations, not every routing system.[15]

Routing is also an attack surface separate from answer generation. Rerouting LLM Routers demonstrated input additions that caused evaluated open-source and commercial routers to select more expensive models. An attacker could thereby affect resource allocation even when answer quality was not degraded. The authors call the broader requirement "LLM control plane integrity": preserving orchestration behavior under adversarial input.[14]

Cost attacks differ from attempts to elicit prohibited content, although both can involve manipulated prompts. Testing only the final answer can miss an unwanted change in model selection. A routing assessment therefore needs to examine its decisions and costs as well as generated text.[14]

Data handling belongs in eligibility rules. Provider selection can change where a request is processed, and the router or embedding service may itself receive prompt data. OpenRouter documents separate controls for provider data collection, zero-data-retention endpoints, and support for requested parameters. These controls illustrate why "choose the cheapest model" is an incomplete policy for requests with handling restrictions.[5][7]

Implementations and operational choices

OpenRouter separates automatic model selection from provider selection. Its Auto Router documentation, reviewed in September 2026, describes task classification followed by selection using recent aggregate spending patterns and configured restrictions. That market-based signal is different from a directly measured probability of factual correctness.[17]

Amazon Bedrock intelligent prompt routing uses predicted response quality and configurable criteria within a selected model family. Its responses identify the model used, allowing applications to observe the actual choice rather than treating the router's identifier as the answering model.[16]

Fallback configuration needs comparable visibility. OpenRouter's model-fallback documentation describes trying alternatives when a selected model encounters an error and returning the model that served the response. A successful fallback is evidence of request completion, not proof that the replacement matched the original model's quality on the task.[18]

A routing system's operating point should therefore be stated as a tested policy with a defined model pool, workload, quality measure, and cost accounting. A provider's supported options explain how to configure a system; workload-specific evaluation establishes whether that configuration achieves its intended result.[16][19]

References

  1. ^1 ^2 ^3 ^4Ong, Isaac, et al. RouteLLM: Learning to Route LLMs with Preference Data. arXiv:2406.18665, 2024; revised February 23, 2025. Published at ICLR 2025 under the title *RouteLLM: Learning to Route LLMs from Preference Data*.
  2. ^1 ^2 ^3 ^4Chen, Lingjiao, Matei Zaharia, and James Zou. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176, May 9, 2023.
  3. ^1 ^2 ^3Hu, Qitian Jason, et al. RouterBench: A Benchmark for Multi-LLM Routing System. arXiv:2403.12031, March 2024.
  4. ^Fedus, William, Barret Zoph, and Noam Shazeer. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. 2021; revised 2022, Journal of Machine Learning Research.
  5. ^1 ^2 ^3 ^4OpenRouter. Provider Routing. Documentation, accessed September 27, 2026.
  6. ^1 ^2 ^3LiteLLM. Router: Load Balancing. Documentation, accessed September 27, 2026.
  7. ^1 ^2 ^3LMSYS. RouteLLM repository and usage documentation. Accessed September 27, 2026.
  8. ^1 ^2 ^3Chen, Shuhao, et al. RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models. NeurIPS 2024.
  9. ^1 ^2Tran, Co, Salman Paracha, Adil Hafeez, and Shuguang Chen. Arch-Router: Aligning LLM Routing with Human Preferences. arXiv:2506.16655, June 19, 2025.
  10. ^1 ^2 ^3Varangot-Reille, Clovis, Christophe Bouvard, and Antoine Gourru. Generalising LLM Routing using Past Performance Retrieval: A Few-Shot Router is Sufficient. EACL 2026 Student Research Workshop, pp. 304-319, March 2026.
  11. ^1 ^2Guo, Chuan, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On Calibration of Modern Neural Networks. ICML 2017, PMLR 70, pp. 1321-1330.
  12. ^1 ^2 ^3 ^4Zheng, Lianmin, et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks Track.
  13. ^1 ^2 ^3Li, Hao, et al. LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. Findings of ACL 2026. Preprint, January 12, 2026.
  14. ^1 ^2Shafran, Avital, Roei Schuster, Thomas Ristenpart, and Vitaly Shmatikov. Rerouting LLM Routers. arXiv:2501.01818, January 3, 2025.
  15. ^1 ^2Kassem, Aly M., Bernhard Schölkopf, and Zhijing Jin. How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities. EACL 2026, pp. 7496-7507. Preprint, 2025.
  16. ^1 ^2 ^3Amazon Web Services. Understanding intelligent prompt routing in Amazon Bedrock. Documentation, accessed September 27, 2026.
  17. ^OpenRouter. Auto Router. Documentation, accessed September 27, 2026.
  18. ^1 ^2 ^3 ^4OpenRouter. Model Fallbacks. Documentation, accessed September 27, 2026.
  19. ^1 ^2 ^3Moslem, Yasmin, and John D. Kelleher. Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey. Transactions on Machine Learning Research, 2026; arXiv:2603.04445, revised August 30, 2026.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 2,511 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent AI-assisted editorial review checked the published text against cited primary documentation and research. Version-specific behavior and study limitations are stated in the article; this is not a guarantee of runtime behavior or factual infallibility.

Cite this page: AI Wiki. "LLM Model Routing." aiwiki.ai, updated 27 Sept 2026, fact-checked 27 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/llm_model_routing

Suggest edit

What links here