Prime Inference
Prime Inference is a model-serving service from Prime Intellect. Its OpenAI-compatible API gives applications access to large language models through a common endpoint and model catalog. The catalog distinguishes models hosted on Prime Intellect's infrastructure from models routed to external providers; the two serving types use the same API key and endpoint.[1]
Prime Intellect announced the product on October 2, 2026, following its September 22 public deployment of GLM-5.3 through OpenRouter.[2] The vLLM project also acknowledged the launch and its use of vLLM for GLM-5.3 serving.[3]
Service and deployment types
| Catalog serving type | Where the model runs | What the API client selects |
|---|---|---|
| Hosted | Prime Intellect's inference infrastructure | The hosted model's catalog ID |
| Gateway | An external inference provider to which Prime routes the request | The gateway model's catalog ID |
At launch, Prime offered serverless endpoints and reserved capacity. It described NVIDIA Blackwell as its current hosted-model hardware and Vera Rubin as a future addition.[2] That hardware description does not establish the infrastructure used by gateway providers.[1]
Availability, prices, request limits, and supported options are model-specific. The Models API returns catalog entries with model IDs, serving types, prices, and specifications when supplied. Clients use the returned id in a completion request; a fixed list of models or prices can become outdated as the catalog changes.[5]
GLM-5.3 deployment
Prime's dedicated model guide identifies z-ai/glm-5.3 as a hosted deployment of GLM-5.3, an open-weights reasoning model from Z.ai. The guide documents these limits and defaults for that endpoint:[4]
| Property | Documented value |
|---|---|
| Model ID | z-ai/glm-5.3 |
| Context window | 1,048,576 tokens |
| Maximum output | 131,072 tokens |
| Input and output | Text |
| Reasoning | Always enabled |
| Reasoning effort | low, high, or max; default max |
| Serving type | Hosted on Prime Intellect GPU infrastructure |
These values describe this model's Prime endpoint. They should not be transferred to another GLM variant, another provider, or every model reachable through Prime Inference.[4]
API access
The API base URL is https://api.pinference.ai/api/v1. Authentication uses a bearer API key with Inference permission.[4][6] Prime's key-management guide supports scoped permissions and expiration dates, recommends setting an expiration, and states that the secret value is displayed only once after creation.[7] The GLM-5.3 guide directs developers to keep the key server-side.[4]
| Operation | Interface | Notes |
|---|---|---|
| Discover models | GET /models | Returns current model IDs and available metadata.[5] |
| Generate a chat response | POST /chat/completions | Accepts a model ID and conversation messages.[6] |
| Stream generated output | Chat request with "stream": true | Uses server-sent events.[6] |
| Request usage details | "usage": {"include": true} | Prime extension for token counts and cost; the Python SDK passes it through extra_body.[6] |
| List models with the CLI | prime inference models | Alternative to the catalog API.[5] |
For example, an HTTP client can submit this request after setting PRIME_API_KEY in its environment:[6]
curl https://api.pinference.ai/api/v1/chat/completions \
-H "Authorization: Bearer $PRIME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-5.3",
"messages": [{"role": "user", "content": "Explain prefix caching in one sentence."}]
}'
The OpenAI Python SDK can use the service by setting base_url to the Prime URL and supplying the Prime key.[1] Supported parameters and limits still depend on the chosen model.[6] OpenAI compatibility here describes the documented API interface; it should not be read as a promise that all OpenAI services, parameters, or model behavior are reproduced.
Billing and usage
Without a team header, an inference request bills the API key owner's personal account. To charge a team, the request must include X-Prime-Team-ID, and the key owner must have access to that team. A team selected in the CLI does not automatically become the billing account for an independent HTTP or SDK request.[8]
For the Python SDK, the overview documents default_headers={"X-Prime-Team-ID": "your-team-id"} on the client; prime config set-team-id controls CLI operations.[1] The team ID can be obtained from the team profile or prime teams list. The insufficient_funds error concerns the balance of the account selected by the key and optional team header, so funding a personal account does not fund a team, or vice versa.[8]
Serving architecture
The launch describes separate prefill and decode GPU pools: NVIDIA Dynamo handles orchestration, vLLM executes the model, and NIXL moves KV cache state. Mooncake supplies a host-DRAM cache tier, and the stack also uses FlashInfer.[2]
In disaggregated serving, prefill processes input tokens while decode generates the continuation. Separating their execution allows different resource allocations and parallelism choices for the two phases. The DistServe paper studies this separation under both time-to-first-token and time-per-output-token requirements, including the cost of transferring state between stages.[10] Those research results explain the design problem; they are not independent measurements of Prime Inference.
The vLLM documentation labels its disaggregated-prefill feature experimental. It describes separate instances connected by a KV-transfer connector, with the aim of independently tuning initial response latency and inter-token latency. It explicitly warns that disaggregation alone does not improve throughput.[12] A deployed service's routing, model, memory capacity, traffic, and interconnect remain part of its performance conditions.
The Mooncake paper describes a cache-centered architecture that separates prefill and decode and uses CPU memory and other cluster storage for cached attention state.[11] The distinction between cached state and output generation matters: vLLM's prefix-caching guide states that reusing a computed prefix saves prefill work but does not speed up the generation of new tokens.[13]
Reported performance
Prime reported GLM-5.3 AgentX tests on GB200 NVL72, mixing returning sessions with cold prompts. Its example turn adds about 6,000 tokens to a 140,000-token prompt. At a 1:4 prefill/decode ratio, it reported 66 sessions per prefill group, 101 end-to-end tokens/s/user, and 100 output tokens/s/GPU.[2]
These are company-reported measurements for that workload and configuration, not a guaranteed rate for every API call or independent verification of a provider ranking. Per-user speed, total output rate per GPU, and the number of concurrent sessions measure different properties. Time to first token also includes work before generation begins, whereas time per output token concerns the generation phase.[10] A comparison that changes prompt lengths, prefix reuse, concurrency, or the latency requirement is not the same experiment.
Adapter serving and roadmap
The adapter guide documents a separate route for existing LoRA adapters from legacy shared-LoRA Hosted Training runs. A completed adapter must have READY status before deployment; after it reaches DEPLOYED, the inference model identifier uses the base_model:adapter_id form. Unloading removes it from serving while preserving its files.[9]
The guide states that the shared training service stops accepting new runs on October 5, 2026, while existing adapters remain downloadable and deployable until further notice. It describes dedicated Hosted Training as a closed beta.[9] This legacy adapter path is narrower than a general promise to deploy any fine-tuned model.
The October 2 launch lists batch/async inference and dedicated one-click deployments of reserved capacity, including fine-tuned models, as roadmap items rather than launch features.[2]
References
- ^1 ^2 ^3 ^4Prime Intellect Docs. "Prime Inference." Hosted and gateway models, API setup, and team billing. Accessed October 5, 2026. docs.primeintellect.ai/...overview
- ^1 ^2 ^3 ^4 ^5Prime Intellect Team. "Prime Inference: Fast, Reliable Serving for Frontier Open Models." October 2, 2026. primeintellect.ai/...prime-inference
- ^vLLM. Announcement acknowledging the Prime Inference launch and GLM-5.3 deployment. X, October 2, 2026. x.com/...2106159655290589305
- ^1 ^2 ^3 ^4Prime Intellect Docs. "GLM-5.3." Prime endpoint limits and reasoning settings. Accessed October 5, 2026. docs.primeintellect.ai/...glm-5-3
- ^1 ^2 ^3Prime Intellect Docs. "Models API." Accessed October 5, 2026. docs.primeintellect.ai/...inference-models
- ^1 ^2 ^3 ^4 ^5 ^6Prime Intellect Docs. "Chat Completions API." Streaming and usage extension. Accessed October 5, 2026. docs.primeintellect.ai/...inference-chat-completions
- ^Prime Intellect Docs. "API keys." Permissions, expiration, and key management. Accessed October 5, 2026. docs.primeintellect.ai/...api-keys
- ^1 ^2Prime Intellect Docs. "Troubleshooting." Inference billing and insufficient funds. Accessed October 5, 2026. docs.primeintellect.ai/...troubleshooting
- ^1 ^2Prime Intellect Docs. "Deploying LoRA Adapters for Inference." Legacy adapter support and Hosted Training transition. Accessed October 5, 2026. docs.primeintellect.ai/...adapter-deployments
- ^1 ^2Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. "DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving." 18th USENIX Symposium on Operating Systems Design and Implementation, July 2024, pp. 193-210. usenix.org/...zhong-yinmin
- ^Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. "Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving." arXiv:2407.00079, first submitted June 24, 2024; version 4, September 3, 2025. arxiv.org/...2407.00079
- ^vLLM documentation. "Disaggregated Prefilling (experimental)." Accessed October 5, 2026. docs.vllm.ai/...disagg_prefill
- ^vLLM documentation. "Automatic Prefix Caching." Accessed October 5, 2026. docs.vllm.ai/...automatic_prefix_caching
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 1,423 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent source review on October 5, 2026. Full new article independently reviewed against the October 2 launch, live hosted/gateway model and API guides, billing and key-management guides, dated legacy LoRA transition, and primary DistServe/Mooncake research. Vendor tests and roadmap status are explicitly qualified.
Cite this page: AI Wiki. "Prime Inference." aiwiki.ai, updated 4 Oct 2026, fact-checked 4 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/prime_inference