Citation and evidence

Decision 2.0

7 min full readUpdated 16 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI InferenceAI ModelsOpen Source AI

Cite this article

Decision 2.0 is a family of open-weight models for structured decisions from the vLLM Semantic Router project. Given text or JSON and a set of questions, the models return probabilities for specified choices, yes/no answers, or ordered rating levels. They do not produce a free-form textual answer. The family is intended for tasks such as classifying a request before selecting a model or applying a routing rule.[2][5][9]

vLLM announced Decision 2.0 on October 3, 2026, describing six size classes from 0.6B to 27B, Apache-2.0 licensing, Transformers loading, and support for answering 64 questions about one request in a forward pass.[1] Decision 2.0 is the model family; vLLM Semantic Router is the routing system that can use it. Neither name denotes a new version of the vLLM inference engine.[2][9]

Released models

The official Hugging Face collection contains six checkpoints. Model names use nominal size labels; the cards' parameter figures differ from those labels for several variants. The table reproduces the parameter and context figures in the October 2026 model cards, with B denoting billions of parameters.[2][3][4][5][6][7][8]

ModelReported parametersContext lengthModel card
Decision-2.0-Kai-0.6B0.60B8,192 tokensKai[3]
Decision-2.0-Eos-0.8B0.75B16,384 tokensEos[4]
Decision-2.0-Sol-2B1.88B16,384 tokensSol[5]
Decision-2.0-Nox-4B4.21B16,384 tokensNox[6]
Decision-2.0-Lux-9B7.94B16,384 tokensLux[7]
Decision-2.0-Vega-27B29.37B32,768 tokensVega[8]

Expanded article table

All six cards identify the release license as Apache-2.0 and document the same three decision types.[3][4][5][6][7][8] The smaller nominal model is not necessarily better or worse on every evaluation: Eos has a higher reported JevArena score than Sol, while Sol has higher reported human-labelled transfer and Jev Decision Index scores.[4][5]

Inputs and decision types

The model-card interface calls system_one(state=..., questions=...). The state supplies the information to examine. The questions mapping gives each question an identifier, a type, instructions, and any required answer criteria. Results appear under answers in the returned object.[5][6]

TypeInput supplied by the applicationMeaning of the result
choiceNamed alternatives with descriptionsProbabilities over alternatives and a selected option
noulA yes/no proposition or questionProbability assigned to the proposition being true
scoreAn ordered list of rating levelsA distribution over the levels; the router can use the expected level

Expanded article table

The router documentation uses these outputs to test predicates. For a Choice condition, it can test the selected label and its probability. Noul conditions test the true probability; Score conditions test the expected level. A probability threshold is a configured policy, not a universal acceptance standard.[10]

For example, a support application could ask for a destination team, whether a receipt is present, and an urgency rating from the same customer record. These are separate predictions about shared evidence. The model's selection does not itself transfer the case, issue a refund, or establish that the underlying record is truthful.[5]

Multiple questions and shared context

The announcement's 64-question example concerns questions about one request, rather than 64 separately generated conversations.[1] The released Nox runtime distinguishes ordinary batched execution from an optional shared-context path. The ordinary path encodes each question with the state and alternatives as its own row, then evaluates the rows as a padded batch. It therefore recomputes the shared input across questions.[11]

With shared context enabled, the runtime can compute a common token prefix once. Its tree mode packs the prefix and question suffixes into one forward operation; a cache mode uses a prefix computation followed by a suffix batch. The implementation can fall back from tree mode when the packed input exceeds its forward-token budget.[11]

The shared-context option is off by default in the inspected native loader. Its source warns that different arithmetic and reduction orders can change close decisions. A configurable margin can send uncertain cases back through the ordinary path. This optimization should not be interpreted as a promise of identical answers under every execution configuration.[11][12]

Use in vLLM Semantic Router

LLM model routing chooses among models or model paths for a request. Decision 2.0 supplies predictions that a configured routing policy can consume; it does not define the entire policy or guarantee the quality of the eventual generated answer.[9]

The router's decision signal sends typed questions to a model-runtime deployment. Questions for the same deployment are grouped into a call. The configuration can combine their results with other conditions, select a model, and decide how to handle an unknown signal. A starting, overloaded, failed, or slow runtime can leave the signal unknown, so error handling is part of the routing configuration.[10]

The documentation recommends these models for semantic judgments such as request type or reasoning difficulty. It distinguishes those tasks from structural checks such as token counts and from specialized signals for personal information or jailbreak detection.[10]

Reported evaluations

The following results are reported by the model publisher, not measurements conducted by AI Wiki. The cards distinguish JevArena, a human-labelled transfer evaluation, and the Jev Decision Index.[3][4][5][6][7][8]

ModelJevArenaHuman-labelled transferJev Decision Index
Kai-0.6B48.645.916.3
Eos-0.8B53.950.320.1
Sol-2B52.151.329.5
Nox-4B63.652.343.8
Lux-9B68.156.246.3
Vega-27B74.058.756.5

Expanded article table

The cards describe JevArena comparisons as using fixed prompts and common scoring, with missing or invalid answers counted as errors. Human-labelled transfer is the median macro-F1 over 15 tasks, multiplied by 100. The Jev Decision Index results use its official 0.2.1 kit on the released weights, while comparison-model figures come from a September 28, 2026 leaderboard snapshot. These metrics have different definitions and should not be read as interchangeable accuracy percentages.[3][4][5][6][7][8]

The Nox and Vega cards qualify their same-size JevArena comparisons: Nox is described as statistically level with Decider 4B, and Vega as statistically level with AutoJev-27B. The reported numerical lead does not establish a significant advantage over those comparators.[6][8]

The cards also report single-question median GPU latencies, but those numbers describe the publisher's measurements rather than a guaranteed service response time. They do not establish the latency of an application that adds networking, queueing, routing, or a downstream model call.[3][7][8]

Loading and interpretation

The documented loading path uses the Transformers AutoModel interface with custom repository code enabled. The cards specify transformers>=5.17, PyTorch and Safetensors; Vega's example also installs PEFT.[3][8] The inspected wrapper loads the packaged runtime and checks its files against a manifest. Vega is distributed as a base-bound adapter package, so loading it also requires the pinned base-model files.[13][16]

Custom model code is executable software. Transformers warns that loading custom models can execute malicious code and recommends pinning a commit revision. Manifest checks identify expected files; they do not independently establish that the code is safe.[13][14]

A returned probability is also not proof of calibration on a deployment's traffic. Calibration concerns whether predicted probabilities correspond to observed outcome frequencies. The ICML 2017 calibration study establishes that distinction for neural classifiers; it did not evaluate Decision 2.0.[15] Applications using scores to trigger consequential actions need evidence on their own task definitions and thresholds, rather than treating a benchmark score or a structured output format as a correctness guarantee.[10][15]

References

  1. ^1 ^2vLLM (@vllm_project). Decision 2.0 release announcement. October 3, 2026.
  2. ^1 ^2 ^3vLLM Semantic Router Team. Decision 2.0 model collection. Accessed October 4, 2026.
  3. ^1 ^2 ^3 ^4 ^5 ^6 ^7vLLM Semantic Router Team. Decision-2.0-Kai-0.6B model card. Accessed October 4, 2026.
  4. ^1 ^2 ^3 ^4 ^5 ^6vLLM Semantic Router Team. Decision-2.0-Eos-0.8B model card. Accessed October 4, 2026.
  5. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9vLLM Semantic Router Team. Decision-2.0-Sol-2B model card. Accessed October 4, 2026.
  6. ^1 ^2 ^3 ^4 ^5 ^6 ^7vLLM Semantic Router Team. Decision-2.0-Nox-4B model card. Accessed October 4, 2026.
  7. ^1 ^2 ^3 ^4 ^5 ^6vLLM Semantic Router Team. Decision-2.0-Lux-9B model card. Accessed October 4, 2026.
  8. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8vLLM Semantic Router Team. Decision-2.0-Vega-27B model card. Accessed October 4, 2026.
  9. ^1 ^2 ^3vLLM Semantic Router Team. Project repository. Accessed October 4, 2026.
  10. ^1 ^2 ^3 ^4vLLM Semantic Router Team. Decision Signal documentation. Accessed October 4, 2026.
  11. ^1 ^2 ^3vLLM Semantic Router Team. Decision 2.0 shared-context implementation. Accessed October 4, 2026.
  12. ^vLLM Semantic Router Team. Decision 2.0 native API implementation. Accessed October 4, 2026.
  13. ^1 ^2vLLM Semantic Router Team. Decision 2.0 Transformers model wrapper. Accessed October 4, 2026.
  14. ^Hugging Face. Loading models: custom models. Accessed October 4, 2026.
  15. ^1 ^2Guo, Chuan, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On Calibration of Modern Neural Networks. ICML 2017, PMLR 70, pp. 1321-1330.
  16. ^vLLM Semantic Router Team. Vega-27B package manifest. Accessed October 4, 2026.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 1,450 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent source review on October 4, 2026. Full draft reviewed. Six numerical model/evaluation rows cross-checked against pinned cards; native runtime opt-in sharing and numerical caveats checked; benchmark claims remain publisher-attributed; generic calibration paper clearly separated from model evaluation.

Cite this page: AI Wiki. "Decision 2.0." aiwiki.ai, updated 3 Oct 2026, fact-checked 3 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/decision_2_0

Suggest edit

What links here