Intern-Decision
Intern-Decision is a family of three open-weight multimodal decision models (Intern-Decision-0.8B, Intern-Decision-2B and Intern-Decision-4B) released in late September 2026 by the InternLM team of Shanghai AI Laboratory. The models do not write free-form text. A caller supplies a "state" (text or structured data, optionally with up to eight images) and a schema of named, typed questions; the model returns an answer and a probability distribution for every question from a single forward pass [1][2]. Each model is a fine-tune of the matching dense Qwen3.5 checkpoint (Qwen3.5-0.8B, 2B and 4B), with the language backbone trained and the vision tower and projector frozen [1][2]. The request and response format deliberately follows the one used by Jev, the hosted decision model that TypeSafe AI launched on 15 September 2026, and the release's headline claim is a comparison with Jev: on the team's own seven-suite accuracy table the 4B model averages 90.02 against Jev's 88.74, with a mean local latency of 44.16 ms against 109.70 ms for Jev [1][2]. All benchmark, latency and calibration figures in this article are self-reported by the Intern-Decision authors unless another source is named.
The weights, the inference wrapper and the model cards are released under the Apache License 2.0, with the upstream Qwen license kept alongside as a separate file [2][14]. The GitHub repository adds training code, two inference backends, evaluation scripts, the test records for the seven accuracy suites, a 96-case calibration benchmark and a browser demo, but not the training data [1].
Overview
| Item | Detail |
|---|---|
| Developer | InternLM team, Shanghai AI Laboratory (repository license copyright "Shanghai AI Laboratory") [1][14] |
| Sizes | 0.8B, 2B, 4B (about 0.85B, 2.21B and 4.54B parameters in the published safetensors) [9] |
| Base models | Qwen/Qwen3.5-0.8B, Qwen/Qwen3.5-2B, Qwen/Qwen3.5-4B [2][3][4] |
| Question types | choice, score, noul (yes/no) [1] |
| Inputs | State plus 1 to 16 questions, up to 62 options per question, up to eight images, 8,192-token default input limit [2] |
| Output | Typed JSON answers with calibrated probabilities, in a Jev-compatible envelope [2] |
| Hugging Face repositories created | 26 September 2026 [9] |
| GitHub repository created | 24 September 2026 (first public code commit 26 September) [10] |
| License | Apache 2.0, with the Qwen license file (LICENSE-QWEN) retained [2][14] |
| Distribution | Hugging Face collection internlm/intern-decision, ModelScope (Shanghai_AI_Laboratory/Intern-Decision-*), Hugging Face demo Space [1][11][12] |
Background
In mid-September 2026 TypeSafe AI released Jev, a hosted model that answers typed questions about a piece of state with probabilities instead of generating text, and published its request format as an HTTP API that requires an API key [15]. Within days, independent developers began publishing open Jev-style models and harnesses. The Intern-Decision benchmark table compares against four of them alongside Jev: Laya, described by its author as a "non-autoregressive System 1 decision engine"; SemIf (formerly OpenJev), whose repository describes it as "semantic ifs from open models, on a 3090 at home" and says it is not affiliated with Jev or TypeSafe; Kev, described as a "Jev-like family of decision models built on top of Qwen3.5/3.8"; and JevK5, billed as an "open-weight alternative to TypeSafe Jev" [1][17][21]. Three of the seven test suites come from JevBench, a benchmark for "Jev-class decision models" run by Benchmark Heaven, which states that it is not affiliated with or endorsed by TypeSafe [17].
Jev itself, according to TypeSafe's Models page, accepts text only: "No image, audio, or video input" [16]. Image input is therefore one of the main functional differences Intern-Decision offers over the system it imitates.
Timeline
| Date (UTC) | Event |
|---|---|
| 24 September 2026 | GitHub repository InternLM/Intern-Decision created [10] |
| 26 September 2026, 05:35 to 05:36 | Hugging Face repositories for the 0.8B, 2B and 4B models created [9] |
| 26 September 2026, 08:03 | First commit to the GitHub repository: "Release Intern-Decision training, inference, demo and evaluation code" [10] |
| 26 September 2026 | Community GGUF conversions of all three sizes appear on Hugging Face [18] |
| 26 September 2026 | Models listed on ModelScope under Shanghai_AI_Laboratory [12] |
| 29 September 2026 | ModelScope's official X account announces the release [13]; Apache 2.0 license file committed to the GitHub repository [10] |
| 29 September 2026 | A pull request proposing native Intern-Decision serving is opened on SGLang [22] |
| 30 September 2026 | A Core ML export of the 0.8B model is published; a request to add Intern-Decision-4B to the JevBench leaderboard is filed [19][23] |
How it works
Structured decisions in one forward pass
The model card calls Intern-Decision a "multimodal structured decision model": it "accepts a shared state, a schema of named questions, and optional images, and returns an answer distribution for every question in one model forward pass" [2]. Inference proceeds in five steps [2]:
- Each question's options are mapped, in their original order, to single-token symbols
AtoZ,atozand0to9. This is why a question can have at most 62 options. - The system prompt, the state, the decision schema and a complete assistant-side JSON skeleton are rendered through the checkpoint's chat template. The skeleton contains one special
<decision>placeholder token per question. - One causal Hugging Face Transformers forward pass is run over the whole prompt.
- For each question, the logits at the position immediately before its placeholder are restricted to that question's legal answer symbols and passed through a softmax; the checkpoint's calibration temperature is then applied.
- The symbols are mapped back to the original option values and returned as typed JSON.
The model card states that this "does not call generate() or sample free-form text" [2]. Because every field is read from the same forward pass, several questions cost roughly one pass rather than one generation each. Question and option order must be preserved [1][2].
Question types and response
The three question types follow Jev's naming [2][15]:
| Type | Criteria | Returned value |
|---|---|---|
choice | Ordered object mapping option values to descriptions | Selected option (choice) plus probabilities over all options |
score | Ordered list (values become "0", "1", ...) or an object with numeric keys | Probability-weighted expected score, which can fall between categories, plus a legend |
noul | Optional descriptions of no and yes | Probability of yes |
Every answer also carries probabilities (the calibrated distribution), confidence (the maximum candidate probability), decision (the highest-probability option, with lexical tie-breaking) and source: local. The response uses the Jev envelope of model, answers and usage, and adds backend, timing and calibration fields as extensions. In usage, output_tokens and decision_count count scored fields rather than generated tokens [2].
A request can contain 1 to 16 questions. Inputs longer than the default limit of 8,192 tokens, including image tokens, are rejected rather than truncated. Images are passed as an ordered list, with up to eight per request; over HTTP they must be sent as uploaded or base64-encoded bytes, not as file paths or remote URLs, with limits of 12 MB per upload, 32 MB combined and 16 million pixels per image [2][6].
ModelScope's announcement summarized the intended uses as "routing, tool selection, scoring, and multimodal workflow control" [13]. The model card's worked example routes a customer message about a double charge: a choice question picks the billing or delivery team, a score question rates urgency, and a noul question asks whether the customer wants a refund, all in one call [2]. The GitHub README shows demonstration videos titled Mario, Mouse Moving, Doom Defend, Doom Health Gathering and Browser Use, and notes that the demo examples are "synthetic interface illustrations" [1].
Training method
According to the GitHub README, the objective is "autoregressive masked language modeling" over the JSON skeleton. Ground-truth answer symbols appear only in the training labels, never in the input, and "the logit immediately before a marker predicts that field's answer", so all fields share one causal pass. Training uses cross-entropy over the full vocabulary; the restriction to legal answer symbols happens only at inference [1]. The data documentation says training uses hard labels, not soft probability targets [6].
The release describes the method, not the data. The data interface document states that it "does not disclose the models' training data sources, composition, quantities, or mixing proportions", and training data, private calibration and validation records, images and preparation pipelines are all excluded from the release [6][1]. The training configuration reads its architecture from the base checkpoint and supports the 0.8B, 2B, 4B and 9B dense Qwen3.5 variants, although only the first three sizes were released [1]. Training runs on InternLM's XTuner framework, pinned to a specific upstream commit [1]. The README and model cards do not link a technical report or paper.
Architecture
The released checkpoints keep the Qwen3.5 architecture (Qwen3_5ForConditionalGeneration) with an added <decision> token. Their configuration files show the hybrid layout of the base models, in which every fourth layer uses full attention and the others use linear attention [9].
| Model | Total parameters (BF16 safetensors) | Language layers | Hidden size | Vision encoder depth |
|---|---|---|---|---|
| Intern-Decision-0.8B | 852,985,920 | 24 | 1,024 | 12 |
| Intern-Decision-2B | 2,213,241,664 | 24 | 2,048 | 24 |
| Intern-Decision-4B | 4,539,265,536 | 32 | 2,560 | 24 |
All three share a 248,320-token vocabulary [9]. The weights are stored as separate language, vision and projector safetensors shards [9].
Calibration
Each checkpoint ships with a single temperature that rescales its candidate probabilities. The model card is explicit that this is "candidate probability calibration, not a sampling temperature": the published formula applies softmax(log(p) / T) to the candidate probabilities, which changes confidence, the noul probability and the expected score but preserves the argmax decision [2]. Per the model cards, each temperature was fitted by minimizing negative log-likelihood on 1,728 designated calibration cases, with 1,693 separate validation cases, and test-suite labels were not used to select it [2].
| Model | Default temperature T |
|---|---|
| Intern-Decision-0.8B | 2.747760550703 |
| Intern-Decision-2B | 2.100509348278 |
| Intern-Decision-4B | 1.99241824 |
Sources: model cards [2][3][4].
The evaluation guide notes that the private data behind these presets is not distributed, "so their fitting experiment cannot be independently reconstructed from this release" [5]. Users fitting their own checkpoints are told to use disjoint calibration data and warned that an NLL-optimal temperature "need not minimize ECE on another dataset" [1]. For background on the metrics, see calibration and expected calibration error.
Benchmark results
Seven-suite accuracy
The accuracy table covers seven suites with 10,751 rows and 12,351 decisions in total [5]. Three are the public portions of JevBench (Easy, Original and Hard); the others are Typed Decision, ToolACE, AG News and WildJailBreak test sets, bundled in the repository with file hashes and upstream attribution [5][8].
| Suite | Rows | Decisions | Upstream source |
|---|---|---|---|
| JevBench-Easy | 48 | 48 | fstandhartinger/jevbench (MIT) |
| JevBench-Original | 72 | 72 | fstandhartinger/jevbench (MIT) |
| JevBench-Hard | 111 | 111 | fstandhartinger/jevbench (MIT) |
| Typed Decision | 400 | 2,000 | LocalLLaMA/typed-decisions |
| ToolACE | 310 | 310 | Team-ACE/ToolACE |
| AG News | 7,600 | 7,600 | fancyzhx/ag_news |
| WildJailBreak | 2,210 | 2,210 | allenai/wildjailbreak |
Sources: evaluation guide and accuracy manifest [5][8].
Results, as reported in the model cards and GitHub README (accuracy in percent; Brier and ECE are calibrated JevBench-Hard metrics, lower is better) [1][2]:
| Model | Easy | Original | Hard | Typed Decision | ToolACE | AG News | WildJailBreak | Average | Brier | ECE |
|---|---|---|---|---|---|---|---|---|---|---|
| Jev | 100.00 | 98.61 | 72.07 | 73.35 | 91.29 | 89.57 | 96.29 | 88.74 | 0.358 | 0.095 |
| Laya | 95.83 | 72.22 | 28.83 | 35.95 | 63.87 | 92.84 | 14.84 | 57.77 | 0.804 | 0.246 |
| SemIf | 100.00 | 98.61 | 61.26 | 62.80 | 85.16 | 89.22 | 92.53 | 84.23 | 0.498 | 0.112 |
| Kev | 100.00 | 93.06 | 45.05 | 65.60 | 87.42 | 89.82 | 75.97 | 79.56 | 0.738 | 0.262 |
| JevK5 | 100.00 | 97.22 | 73.87 | 64.50 | 80.97 | 89.13 | 90.45 | 85.16 | 0.366 | 0.047 |
| Intern-Decision-0.8B | 97.92 | 80.56 | 52.25 | 77.35 | 94.52 | 88.61 | 64.48 | 79.38 | 0.530 | 0.066 |
| Intern-Decision-2B | 100.00 | 84.72 | 63.96 | 79.35 | 96.45 | 89.96 | 78.33 | 84.68 | 0.437 | 0.100 |
| Intern-Decision-4B | 100.00 | 98.61 | 73.87 | 80.55 | 96.45 | 90.82 | 89.86 | 90.02 | 0.347 | 0.065 |
Several details of the table matter for reading it:
- The average is unweighted. It is the arithmetic mean of the seven suite accuracies, so the 48-item JevBench-Easy counts as much as the 7,600-row AG News [1][5].
- Where the 4B leads and trails Jev. Intern-Decision-4B scores higher than Jev on Typed Decision (80.55 vs 73.35), ToolACE (96.45 vs 91.29), AG News (90.82 vs 89.57) and JevBench-Hard (73.87 vs 72.07), ties it on Easy and Original, and trails it on WildJailBreak (89.86 vs 96.29) [2].
- Calibration on JevBench-Hard. The 4B has the lowest Brier score in the table (0.347 vs Jev's 0.358) and a lower ECE than Jev (0.065 vs 0.095), consistent with ModelScope's claim that it achieves "better probability calibration" [2][13]. JevK5 has the lowest ECE in the table (0.047), and the 2B's ECE (0.100) is slightly higher than Jev's [2].
- Typed Decision is scored per decision, not per row: a row with five fields contributes five decisions [6].
- Backends and temperatures. The Intern-Decision rows were produced with the XTuner backend, with each model's fixed temperature applied afterward; baseline probabilities are their reported or default values, with no extra temperature fitted [1][5].
What "JevBench" means here
The three JevBench suites in the table are the public item sets from Benchmark Heaven's JevBench repository, which also holds back private items: its hard tier has 220 decisions, of which 111 are public, and its easy tier has 72, of which 48 are public [17]. The official JevBench score is a composite of intelligence, calibration, speed and cost measured partly on sealed items, so the per-suite accuracies in the Intern-Decision table are not JevBench leaderboard scores [17]. The Intern-Decision evaluation guide itself says that "public Jevbench is a development diagnostic, not independent held-out evidence or the complete leaderboard" [5]. On 30 September 2026 a GitHub account that had committed to the Intern-Decision repository filed a request for Intern-Decision-4B to be evaluated under JevBench's v1.5 protocol as an offline artifact; as of that date the request was open [23][10].
Latency
Latency was measured on one NVIDIA RTX 4090 with BF16 weights, SDPA attention, native Hugging Face inference, Transformers 5.14.1 and flash-linear-attention 0.4.2. Each request had 289 input tokens and three fields (choice, yes/no and score) answered in one forward pass with the checkpoint's temperature applied [1].
| Model | Mean | Median (P50) | P95 |
|---|---|---|---|
| Jev | 109.70 ms | 106.30 ms | 146.70 ms |
| Intern-Decision-0.8B | 33.98 ms | 33.44 ms | 37.50 ms |
| Intern-Decision-2B | 33.28 ms | 33.15 ms | 33.55 ms |
| Intern-Decision-4B | 44.16 ms | 44.03 ms | 44.60 ms |
Sources: GitHub README and model cards [1][2].
The model card describes these as "per-query end-to-end latency" that is "workload and hardware dependent" [2], and the README warns that "the protocols and scope differ" across the accuracy, latency and calibration comparisons [1]. ModelScope's post put the Jev figure "in the same local HF setup" [13]. Neither the README nor the model cards explain how Jev was timed on that setup; Jev is a hosted model that callers reach through TypeSafe's HTTP API with an API key [15], so its figure cannot be a local forward pass of the same kind. In the reported numbers the 2B model's mean latency is slightly lower than the 0.8B's [1].
Known-distribution calibration pilot
The repository also includes a 96-case "known-distribution" calibration benchmark. It is synthetic and evaluation-only, with six categories, 24 families and 48 paired parameter settings; each question asks about the outcome of a random process (fair dice, socks drawn without replacement, informed and uninformed Monty Hall hosts, birthday collisions and similar) whose exact outcome distribution is known, and predictions are scored against those exact distributions rather than sampled labels [7]. The authors call it "a small diagnostic study, not 96 independent trials" [1].
Each cell below is expected multiclass Brier score / expected ECE (lower is better) [1][2]:
| Category | 4B, uncalibrated | 4B, calibrated (T=1.992418) | Jev (jev-1.13.0) |
|---|---|---|---|
| Direct randomness and support | 0.483 / 0.181 | 0.421 / 0.129 | 0.490 / 0.216 |
| Composed events and mixtures | 0.677 / 0.254 | 0.577 / 0.150 | 0.682 / 0.274 |
| History, conditioning, and hidden state | 0.711 / 0.219 | 0.613 / 0.108 | 0.657 / 0.113 |
| Daily evidence and observation bias | 0.701 / 0.328 | 0.575 / 0.210 | 0.603 / 0.114 |
| Selective disclosure and probability puzzles | 0.540 / 0.119 | 0.510 / 0.049 | 0.483 / 0.138 |
| Sequential and combinatorial processes | 0.656 / 0.180 | 0.605 / 0.058 | 0.657 / 0.116 |
| Overall (pooled) | 0.628 / 0.213 | 0.550 / 0.089 | 0.595 / 0.130 |
All three settings produced 96 of 96 valid predictions. Calibration lowered the 4B's overall Brier score from 0.628 to 0.550 and its ECE from 0.213 to 0.089, below Jev's 0.595 and 0.130; the pilot was not used to fit or select the temperature [1]. Before calibration the 4B was worse than Jev on both overall metrics, and even after calibration Jev has a lower ECE on "daily evidence and observation bias" (0.114 vs 0.210) and a lower Brier score on "selective disclosure and probability puzzles" (0.483 vs 0.510) [1]. The Jev results used version jev-1.13.0 [1], which TypeSafe's Models page still listed as the target of the jev-latest alias at the end of September 2026 [16]. The README does not name the Jev version used for the accuracy and latency tables.
Usage
Model-card wrapper
Each Hugging Face repository ships an inference.py with a DecisionEngine class. After installing the repository's requirements.txt (Python 3.12 or later), a caller creates the engine once and passes one request dictionary per call to predict(), which returns a JSON-serializable, Jev-compatible response [2]. The wrapper supports only the Hugging Face backend, uses the temperature shipped with that size by default, and accepts temperature=1 for uncalibrated probabilities [2]. Image paths in a request are resolved against a configurable media root [2].
Repository backends and HTTP service
The GitHub repository provides two inference backends: native Hugging Face inference (loading Qwen3_5ForConditionalGeneration directly) and InternLM's XTuner, installed in separate virtual environments because they pin different Transformers versions (5.14.1 and 4.57.0) [1]. The evaluation guide notes that the reported accuracy rows used XTuner and that kernel and BF16 differences between backends "can change probabilities and occasionally labels" [5].
A browser demo and HTTP API expose POST /v1/decisions (with the alias /v1/jev), GET /health and interactive documentation. The service binds to loopback by default and does not log request content; the README advises adding authentication and request-size limits before exposing it publicly [1]. An optional "thinking handoff" to an external model, configured with the user's own endpoint and key, is disabled by default; its outputs are not calibrated local probabilities and were not used in the reported evaluations [1].
Hosted demo
The official Hugging Face Space internlm/Intern-Decision runs Intern-Decision-0.8B on CPU in a Docker container, with the 0.8B temperature applied. It accepts typed questions and image uploads, displays probability bars and timing, and exposes the same /v1/jev and /v1/decisions endpoints. Its README notes that CPU predictions "have not been checked for numerical parity with the GPU evaluation results" [11].
License
The Hugging Face repositories are tagged Apache 2.0 and contain two license files: LICENSE, the Apache 2.0 text with a Shanghai AI Laboratory copyright notice, and LICENSE-QWEN, the Apache 2.0 license from Alibaba Cloud's Qwen release [14][9]. The model card says Intern-Decision "is derived from the Qwen3.5 series", asks redistributors to "retain the license and applicable upstream notices", and describes the weights as "modified by decision tuning" [2]. The GitHub repository's code is Apache 2.0; its README adds that "Qwen model terms apply to weights", XTuner is Apache 2.0 and JevBench is MIT [1]. The bundled test sets keep their own terms: the evaluation guide notes that AG News's source metadata reports an unknown license and tells users to check upstream terms before redistribution [5].
Community ports and integrations
As of 30 September 2026, third parties had published several conversions and integrations. All are unofficial.
| Project | Date (UTC) | What it is |
|---|---|---|
| bombdefuser-124/Intern-Decision-{0.8B,2B,4B}-GGUF | 26 September 2026 | GGUF conversions for llama.cpp (FP16 and Q8_0 language models plus an FP16 vision projector in each repository). The 4B card notes that the initial 4B conversion metadata declared 33 blocks instead of 32 and was corrected, and warns that Q8_0 quantization "can change calibration quality" [18] |
| dex0shubham/intern-decision-mlx | 27 September 2026 | Runs Intern-Decision-0.8B on Apple silicon through mlx-vlm and adds the typed readout, calibration and a local /v1/systemone endpoint. Its author reports about 0.9 seconds per decision on a 1080p screenshot on an 8 GB M2 MacBook Air [20] |
| SGLang pull request #41712 | 29 September 2026 (open) | Proposes serving decision-model checkpoints natively in SGLang on /v1/decisions, /v1/jev and /v1/systemone, with Intern-Decision as the first supported family. The author reports Typed Decision accuracy of 77.60% (0.8B) and 80.50% (4B) on one H200, against the official 77.35% and 80.55% [22] |
| FluidInference/intern-decision-0.8b-coreml | 30 September 2026 | Core ML export of the 0.8B text path (no vision tower) for iOS 17 and macOS 14 and later, in fixed-length packages up to 1,024 tokens. Its fidelity table reports no changed top answers for the fp16 packages and 2 of 240 changed for an int8 variant [19] |
Limitations
The authors' own documentation lists several limits:
- Self-reported, partly non-reproducible comparisons. The release regenerates the Intern-Decision rows but "does not include those baseline runners or historical responses", so the external baseline rows cannot be rebuilt from the release alone, and the temperature-fitting data is private [5].
- Public benchmark items. The JevBench items used are public and can be trained on or selected against; the guide calls them a development diagnostic [5][17].
- Small calibration pilot. The 96-case pilot has 48 paired settings, not independent trials [1].
- Backend sensitivity. Different kernels and BF16 rounding "can produce probability differences between backends", and the authors say not to claim bitwise equivalence [1].
- Hard limits on input. At most 62 options per question, 16 questions and eight images per request, and an 8,192-token default input budget with rejection rather than truncation [2].
- No text generation. The model scores candidates; it does not produce explanations or free text [2].
- Undisclosed training data. The sources, composition and size of the training data are not described [6].
See also
- Jev
- TypeSafe AI
- Structured output
- Calibration (machine learning)
- Qwen3.5
- InternVL
- ModelScope
- Small language model
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33InternLM, "Intern-Decision" (GitHub README), InternLM/Intern-Decision repository, accessed 30 September 2026. github.com/...Intern-Decision
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32InternLM, "Intern-Decision-4B" model card, Hugging Face, accessed 30 September 2026. huggingface.co/...Intern-Decision-4B
- ^1 ^2InternLM, "Intern-Decision-0.8B" model card, Hugging Face, accessed 30 September 2026. huggingface.co/...Intern-Decision-0.8B
- ^1 ^2InternLM, "Intern-Decision-2B" model card, Hugging Face, accessed 30 September 2026. huggingface.co/...Intern-Decision-2B
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11InternLM, "Reproducing evaluation tables" (docs/EVALUATION.md), Intern-Decision repository, accessed 30 September 2026. github.com/...EVALUATION.md
- ^1 ^2 ^3 ^4 ^5InternLM, "Data interface" (docs/DATA.md), Intern-Decision repository, accessed 30 September 2026. github.com/...DATA.md
- ^InternLM, "Known-distribution calibration benchmark" (docs/CALIBRATION_BENCHMARK.md), Intern-Decision repository, accessed 30 September 2026. github.com/...CALIBRATION_BENCHMARK.md
- ^1 ^2InternLM, accuracy-v1 benchmark manifest (benchmarks/accuracy-v1/manifest.json), Intern-Decision repository, accessed 30 September 2026. github.com/...manifest.json
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Hugging Face Hub, repository metadata and config.json for internlm/Intern-Decision-0.8B, -2B and -4B (creation times, safetensors parameter counts, file lists), accessed 30 September 2026. huggingface.co/...Intern-Decision-4B
- ^1 ^2 ^3 ^4 ^5GitHub, InternLM/Intern-Decision repository metadata and commit history, accessed 30 September 2026. github.com/...main
- ^1 ^2InternLM, "Intern-Decision" Hugging Face Space README, accessed 30 September 2026. huggingface.co/...Intern-Decision
- ^1 ^2ModelScope, "Shanghai_AI_Laboratory/Intern-Decision-4B" model page and API record, accessed 30 September 2026. modelscope.cn/...Intern-Decision-4B
- ^1 ^2 ^3 ^4ModelScope (@ModelScope2022), post announcing Intern-Decision, X, 29 September 2026. x.com/...2104836200590606542
- ^1 ^2 ^3 ^4InternLM, LICENSE and LICENSE-QWEN files, internlm/Intern-Decision-4B on Hugging Face, accessed 30 September 2026. huggingface.co/...LICENSE-QWEN
- ^1 ^2 ^3TypeSafe AI, "API reference", TypeSafe documentation, accessed 30 September 2026. docs.typesafe.ai/api
- ^1 ^2TypeSafe AI, "Models", TypeSafe documentation, accessed 30 September 2026. docs.typesafe.ai/models
- ^1 ^2 ^3 ^4 ^5Benchmark Heaven, "JevBench" README, fstandhartinger/jevbench repository, accessed 30 September 2026. github.com/...jevbench
- ^1 ^2bombdefuser-124, "Intern-Decision-4B GGUF" model card, Hugging Face, accessed 30 September 2026. huggingface.co/...Intern-Decision-4B-GGUF
- ^1 ^2FluidInference, "Intern-Decision-0.8B for Core ML" model card, Hugging Face, accessed 30 September 2026. huggingface.co/...intern-decision-0.8b-coreml
- ^dex0shubham, "intern-decision-mlx" README, GitHub, accessed 30 September 2026. github.com/...intern-decision-mlx
- ^Repository descriptions of the baseline projects, GitHub, accessed 30 September 2026: NandhaKishorM/laya (github.com/...laya), TheoLeeCJ/SemIf-OpenJev (github.com/...SemIf-OpenJev), jaredpalmer/kev (github.com/...kev), allebee/jevk5 (github.com/...jevk5).
- ^1 ^2"Serve decision model checkpoints natively on /v1/decisions, /v1/jev, and /v1/systemone", sgl-project/sglang pull request #41712, GitHub, opened 29 September 2026. github.com/...41712
- ^1 ^2"[bench request]: Intern-Decision-4B (offline Hugging Face artifact, native TypeSafe probabilities)", fstandhartinger/jevbench issue #158, GitHub, opened 30 September 2026. github.com/...158
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,260 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification (xg15 V5, 30 Sep 2026): ~225 claims across both pages vs GitHub README/docs, HF cards and APIs, config.json, SGLang/JevBench PRs; 2 minor defects fixed in v2
Cite this page: AI Wiki. "Intern-Decision." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/intern_decision