Grok 4.5

RawGraph

Last edited

Fact-checked

In review queue

Sources

9 citations

Revision

v1 · 1,453 words

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Grok 4.5 is a proprietary multimodal large language model and reasoning model in the Grok family. It was developed by SpaceXAI in collaboration with Cursor and released through the xAI API on July 8, 2026. The model accepts text and images, produces text, and has a 500,000-token context window. Its launch emphasized software engineering, tool-using agents, and professional knowledge work rather than a general consumer-chat release.[1][3][4]

Release and positioning

SpaceXAI's release notes place the initial API release on July 8, 2026. Cursor published a joint-release post the same day, and Reuters reported that access began immediately through the API, Grok Build, and Cursor.[1][7][9] The current SpaceXAI announcement page displays July 16 and describes the model as launching that day, but the official release notes separately document both the July 8 API launch and July 17 availability for users of the EU API console.[1][2] On that evidence, July 8 is the initial release date, while the later date on the announcement page reflects a subsequent publication or rollout record whose cause SpaceXAI has not explained.

Grok 4.5 was positioned as a model for coding, agentic tasks, and knowledge work. Cursor made it available in its desktop, web, iOS, command-line, and software-development-kit products at launch. SpaceXAI offered it through Grok Build and its developer platform, while the model card also listed Microsoft Office add-ins and third-party API gateways. The model card said availability on consumer platforms, including the Grok website, mobile applications, and X, was planned for later.[3][6][7]

Model and training

SpaceXAI describes Grok 4.5 as a proprietary text model with image understanding. It accepts text and JPEG or PNG images, but its documented output modality is text rather than generated images. The API's knowledge cutoff is February 1, 2026, and the model catalog says it has no real-time knowledge unless a search tool is enabled.[3][4]

The model card says its training data included public data, internally generated material, and other licensed or rights-secured sources. Post-training used supervised fine-tuning and reinforcement learning with human and synthetic reward signals. It also received supplemental training on anonymized Cursor workflow data intended to improve coding and agentic performance.[6]

SpaceXAI reported that training ran across tens of thousands of Nvidia GB300 graphics processors. Its announcement describes data deduplication, quality scoring, and domain-focused selection, followed by reinforcement learning on hundreds of thousands of tasks. These tasks centered on multistep software engineering and technical work, using automated and model-based grading and long-running asynchronous agent rollouts.[2] The company has not disclosed the model's parameter count or enough architectural detail to reproduce it.

PropertyDocumented value
DevelopersSpaceXAI and Cursor
Initial releaseJuly 8, 2026
Model identifiersgrok-4.5, grok-4.5-latest
Input and outputText and image input; text output
Context window500,000 tokens
Knowledge cutoffFebruary 1, 2026
Reasoning controlLow, medium, or high effort; high by default
Access at launchxAI API, Grok Build, and Cursor
LicenseProprietary

Capabilities and API

The model works with both the xAI Responses API and Chat Completions API. It supports client-defined function calling and SpaceXAI's server-side web search, X search, and code-execution tools. These interfaces allow an application to combine the model's generated text with external information or executable actions, but they do not make the underlying model's stored knowledge current.[3][4]

Reasoning effort is configurable as low, medium, or high, with high used by default. This setting lets developers trade latency and token use against additional inference-time reasoning. Grok 4.5 also supports image inputs for tasks such as interpreting screenshots or visual documents, while professional integrations described by SpaceXAI include drafting and editing work in Word, PowerPoint, and Excel.[2][3]

The model's principal advertised capability is long-horizon software engineering. SpaceXAI and Cursor describe it as able to inspect codebases, edit multiple files, use terminals, run tests, and iterate on failures. Those operations depend on an agent host that supplies the relevant tools and permissions. The distinction matters: the model generates decisions and tool calls, while Grok Build, Cursor, or another application controls execution.[2][7]

Access and pricing

The API model ID is grok-4.5; grok-4.5-latest tracks the current version. Standard pricing for requests with fewer than 200,000 input tokens is $2 per million input tokens, $0.30 per million cached input tokens, and $6 per million output tokens. When a request contains 200,000 or more input tokens, the respective prices double to $4, $0.60, and $12.[4][5]

Server-side tool use is billed separately from tokens. SpaceXAI lists web search, X search, and code execution at $5 per 1,000 successful invocations, collections search at $2.50 per 1,000, and attachment search at $10 per 1,000. Consequently, the advertised $2 and $6 token rates do not represent the full cost of every agentic workload.[5]

SpaceXAI also distributed the model through several gateways, including OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic. Availability was geographically staggered: the company recorded EU API-console access on July 17, nine days after the initial API release.[1][3]

Evaluation

SpaceXAI published results on several software-engineering benchmarks. The following scores are developer-reported and were produced with different benchmark harnesses, so they are not interchangeable measures of one general capability.[2][6]

EvaluationGrok 4.5 resultWhat it measures
DeepSWE 1.062.0%Resolution of software-engineering tasks
DeepSWE 1.153.0%Software tasks run with the mini-swe-agent harness
SWE Marathon29.0%Resolution of longer software tasks
Terminal-Bench 2.183.3%Agent performance in terminal environments
SWE-Bench Pro64.7%Resolution of professional software issues
SWE-Bench Multilingual78.0%Software issues across programming languages

These figures show strong performance on the selected engineering tasks, but they do not establish that Grok 4.5 leads every competing model. In SpaceXAI's own launch chart, another evaluated model scored higher on each of the five prominently displayed benchmarks. Results also depend on tool configuration, inference budget, task selection, and harness design.[2]

Artificial Analysis provides an independent, broader evaluation. As accessed on July 24, 2026, its live model page assigned Grok 4.5 an Intelligence Index score of 54. The index combined tests spanning professional work, banking agents, terminal tasks, science, factual knowledge, and long-context reasoning. The service measured 69.2 output tokens per second and a 10.71-second time to the first answer token. These are dated observations from one provider and can change as the evaluator revises its suite or serving configuration.[8]

One disclosed contamination issue limits interpretation of a separate coding result. Cursor said an earlier snapshot of its own codebase had accidentally entered Grok 4.5's training data, creating an advantage on CursorBench. Because the exact effect could not be measured, Cursor removed CursorBench from its release comparison and said the data would be excluded from future model training. The disclosure does not invalidate unrelated evaluations, but it illustrates why training overlap must be considered when interpreting coding benchmark scores.[7]

Safety and limitations

The model card reports testing for disallowed content, jailbreaks, cybersecurity behavior, and chemical, biological, radiological, and nuclear risks. It characterizes the model as below the company's threshold for dangerous biological and chemical assistance, while also documenting substantial cybersecurity capability. These are developer-run evaluations rather than independent safety audits.[6]

SpaceXAI states that Grok 4.5 is not intended to make autonomous high-stakes decisions in medicine, law, finance, or safety-critical systems without human oversight and validation by domain experts. More generally, its long context window and access to tools do not remove familiar language-model failure modes such as incorrect factual claims, flawed code, missed constraints, or inappropriate actions. Applications using terminal or search access therefore need permission boundaries, output review, and task-specific testing.[4][6]

The model is also closed rather than openly reproducible. Its weights, parameter count, complete data mixture, and full training process are not public. Evaluation claims should therefore be read together with the model card, the disclosed CursorBench overlap, and independent measurements rather than treated as a complete account of performance.[6][7][8]

References

  1. SpaceXAI. "Release Notes." SpaceXAI Docs. Updated July 23, 2026. https://docs.x.ai/developers/release-notes
  2. SpaceXAI. "Introducing Grok 4.5." July 16, 2026. https://x.ai/news/grok-4-5
  3. SpaceXAI. "Grok 4.5." SpaceXAI Docs. Updated July 17, 2026. https://docs.x.ai/developers/grok-4-5
  4. SpaceXAI. "Models and Pricing." SpaceXAI Docs. Updated July 9, 2026. https://docs.x.ai/developers/models
  5. SpaceXAI. "Pricing." SpaceXAI Docs. Accessed July 24, 2026. https://docs.x.ai/developers/pricing
  6. SpaceXAI. "Model Card: Grok 4.5." July 14, 2026, revised July 20, 2026. https://media.x.ai/v1/website/4p5-5184fdf9.pdf
  7. Cursor. "Introducing Grok 4.5." July 8, 2026. https://cursor.com/blog/grok-4-5
  8. Artificial Analysis. "Grok 4.5 (high): Intelligence, Performance and Price Analysis." Accessed July 24, 2026. https://artificialanalysis.ai/models/grok-4-5/
  9. Reuters. "SpaceXAI launches Grok 4.5 model for coding, agentic tasks." July 8, 2026. https://www.investing.com/news/stock-market-news/spacexai-launches-grok-45-model-for-coding-agentic-tasks-4782511

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

Suggest edit