Grok 4.7
| Field | Value |
|---|---|
| Developer | SpaceXAI |
| Release date | September 21, 2026 |
| Model identifier | grok-4.7 |
| Predecessor | Grok 4.6 |
| Architecture | Not formally published; "a new, larger base model" than Grok 4.6 per SpaceXAI; 2.1 trillion parameters per Elon Musk |
| Input and output | Text and image input; text output |
| Context window | 500,000 tokens |
| Knowledge cutoff | May 2026 |
| Reasoning control | Low, medium, high (default), or xhigh; cannot be disabled |
| Pricing | $2 per million input tokens, $6 per million output tokens (doubled at 200,000+ prompt tokens) |
| Availability at launch | xAI API, Grok Build, Cursor, GitHub Copilot, model gateways |
| License | Proprietary |
Grok 4.7 is a proprietary large language model and reasoning model in the Grok family, released by SpaceXAI on September 21, 2026. SpaceXAI calls it its "most capable model for coding and knowledge work" and says it "works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date." It uses a new, larger base model than Grok 4.6 and keeps that model's $2 and $6 per-million-token prices and 500,000-token context window.[1][2][3]
The release came after a run of public schedule changes by Elon Musk, who had first said in July 2026 that Grok 4.7 would follow Grok 4.6 by about two weeks.[21][22][24][25] On Artificial Analysis' Intelligence Index v4.3.2, Grok 4.7 at xhigh reasoning effort scored 46, two points above Grok 4.6 (high) and about six to seven points below Claude Fable 5.1 and GPT-6 Astra, which led the index on launch day. The evaluator also found that the model used far more output tokens per task than its predecessor.[16][17][18][29]
Release
Schedule before launch
Musk gave dates for Grok 4.7 several times on X and on an earnings call before it shipped.
| Date (UTC) | Statement | Implied release |
|---|---|---|
| July 24, 2026 | "Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks," in a reply to his own post on X.[21] | Around August 21 |
| July 28, 2026 | "Grok 4.7 will be the 2.1T model released a few weeks later," after Grok 4.6, which he expected "around August 7." He added that it "will be better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency."[22] | Late August |
| August 4, 2026 | On SpaceX's first earnings call: "We have Grok 4.6 coming out next week, then Grok 4.7 about 3-4 weeks from today."[28] | Late August to early September |
| August 12, 2026 | "Grok 4.7 will exceed all current models," in a reply posted on Grok 4.6's launch day.[23] | Not stated |
| September 2, 2026 | "Grok 4.7 comes out in 10 days."[24] | Around September 12 |
| September 11, 2026 | "Grok 4.7 needs a few more days to cook." He said reinforcement learning "might have penalized response length too much (or something)," so that the model "still gives up on hard tasks (that it can do!) too early and isn't yet sufficiently rigorous in checking its work."[25] | Not stated |
| September 14, 2026 | "Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance."[26] | Not stated |
The model shipped on September 21, a month after the date implied by Musk's July 24 post and nine days after the "10 days" estimate of September 2.[3][21][24] Musk's expectations also came down over the period. On August 12 he said the model would exceed every model then available. On September 14 he wrote that it would be "roughly on par with Opus 5.0, not 5.1," a comparison to Claude Opus 5.[23][26] The August 12 post was no longer on X when checked on September 23, 2026; X showed it as deleted by its author.[23]
The launch post's two headline claims match the two faults Musk had described on September 11. He had said the model gave up on hard tasks too early and did not check its work carefully enough. SpaceXAI's announcement says Grok 4.7 "works longer on difficult tasks" and "checks its own work more carefully."[1][25]
Launch
SpaceXAI's developer release notes record the model's arrival on the xAI API on September 21, 2026 under the identifier grok-4.7.[3] The @SpaceXAI account announced it at 16:17 UTC that day in a thread that opened "Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed." The thread also carried a CursorBench cost chart, a side-by-side video of Grok 4.7 and Grok 4.6 building "an open world city game," and a note that the model was "available now in Cursor and Grok Build, and the Grok API."[9][10][32][11] Eight minutes later, Musk quote-posted the thread's second post and called Grok 4.7 "a strong combination of intelligence, speed & low cost."[15]
The announcement lists Cursor and Grok Build first, then the Grok API, "third-party coding harnesses, and model routers and cloud platforms."[1] SpaceXAI's model overview names the OpenRouter, Vercel and Cloudflare gateways, says Grok 4.7 is the default model in the Grok Build coding agent, and says it is available on all Cursor plans.[2] OpenRouter created its listing at 16:19 UTC on launch day with a 500,000-token context and a dated snapshot name of grok-4.7-20260916.[14] Cursor published a short post, under the Cursor Team byline, calling Grok 4.7 "our most capable model for long-running coding and knowledge work" and linking to SpaceXAI's announcement.[12] Grok 4.6 had been released jointly with Cursor after SpaceX agreed to acquire Cursor's parent company, Anysphere; that background is covered at Grok 4.5.
GitHub began a gradual rollout of Grok 4.7 in GitHub Copilot the same day, for the Copilot Pro, Pro+, Max, Business and Enterprise plans. The model appears in the model picker in Visual Studio Code, Visual Studio, Copilot CLI, the Copilot cloud agent, the GitHub Copilot app, JetBrains IDEs, Xcode and Eclipse, and is billed at provider list prices under usage-based billing.[13]
Model and training
The launch post, model documentation and release notes do not give an architecture description or parameter count, and do not link to a model card. The launch post says only that the model "uses a new, larger base model compared to Grok 4.6." The 2.1 trillion parameter figure comes from Musk's July 28 post. He attached 1.5 trillion parameters to Grok 4.6 in the same post.[1][22]
According to the announcement, Grok 4.7 "was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete." SpaceXAI says this made the model better at verifying its own work and managing longer context. It also says it trained Grok 4.7 "to natively understand the Grok Bot harness, making it better at conversational tasks and general knowledge work."[1] Grok Bot is SpaceXAI's product for persistent AI agents that run on a cloud computer.[3] The launch post does not say how the self-verification or long-context gains were achieved or measured. The New Stack noted that the company "did not disclose whether the context gains came from architectural changes, summarization, retrieval, or better retention across long sequences."[30]
On July 21, Musk said SpaceX's "massive corpus of world-class engineering data (excluding material blocked by ITAR) will be added during supplemental training of the 2T run."[33] The Grok 4.7 announcement does not mention SpaceX data.[1]
API and features
The documented interface is close to Grok 4.6's: text and JPEG or PNG image input, text output, a 500,000-token context window, no text output limit, and support for the Responses and Chat Completions APIs, function calling, structured outputs, and SpaceXAI's server-side web search, X search and code-execution tools. The knowledge cutoff is May 2026. As with earlier Grok models, the model has no access to current events unless a search tool is enabled.[2][4][8]
Reasoning effort can be set to low, medium, high (the default) or xhigh, and reasoning cannot be turned off.[6] One behavior is new for Grok 4.7: on the Responses API, the model always returns reasoning.encrypted_content, even when a request does not ask for it, so a client can pass the encrypted reasoning back in the next turn. SpaceXAI's reasoning documentation says this default "does not stop xAI from storing the thinking trace server-side for rehydration." SpaceXAI's reasoning documentation also says it exposes summaries of grok-4.7's internal reasoning, which can be streamed alongside the answer.[2][6]
Grok 4.7 is one of two models (with Grok 4.6) served on SpaceXAI's US regional endpoint, https://us.api.x.ai/v1. SpaceXAI says requests to that endpoint keep API request handling, inference, moderation and retained request data in the United States; the guarantee does not cover Files, Collections or server-side tools.[7] The model's console page lists a default rate limit of 150 requests per second and 50 million tokens per minute, and states that the Batch API is not supported.[8] As with Grok 4.6, SpaceXAI recommends setting a prompt_cache_key so that a conversation's requests reach the same server. It warns that without one "you often pay full input price on a cache-cold server."[2]
Pricing
| Tier | Input (per 1M tokens) | Cached input | Output |
|---|---|---|---|
| Standard, under 200,000 prompt tokens | $2.00 | $0.50 | $6.00 |
| Standard, 200,000+ prompt tokens | $4.00 | $1.00 | $12.00 |
| US regional endpoint, under 200,000 | $2.20 | $0.55 | $6.60 |
| US regional endpoint, 200,000+ | $4.40 | $1.10 | $13.20 |
| Grok 4.7 Fast, below 200,000 | $4.00 | $1.00 | $12.00 |
| Grok 4.7 Fast, above 200,000 | $6.00 | $1.50 | $18.00 |
Source: SpaceXAI pricing documentation, accessed September 23, 2026.[5]
The standard rates are identical to Grok 4.6's, including the $0.50 cached-input rate. Requests whose prompt reaches 200,000 tokens are billed at the higher rate for every token in the request.[4][5] The US endpoint adds a 10% premium.[5][7]
Grok 4.7 Fast is, in SpaceXAI's description, "the same Grok 4.7 model served on faster infrastructure, at twice the standard token rates." It is sold only through Cursor and Grok Build, is not on the public xAI API, and is not included in Grok Build's free tier. The announcement describes it as having "twice the output speed at twice the price."[1][2][5] The published Fast rates double the standard rates below 200,000 prompt tokens. Above that threshold they are 1.5 times the standard long-context rates.[5]
Benchmarks
SpaceXAI's comparison table
The launch post compares Grok 4.7 at xhigh reasoning effort against Grok 4.6 at high, OpenAI's GPT-5.6 Sol at its maximum setting, and Anthropic's Claude Fable 5.1 at its maximum setting. These are SpaceXAI's figures, and the post does not say how the competitor scores were obtained. Bold marks the highest score in each row.[1][9]
| Evaluation | Grok 4.7 (xhigh) | Grok 4.6 (high) | GPT-5.6 Sol (max) | Claude Fable 5.1 (max) |
|---|---|---|---|---|
| Input price, $ per million tokens | $2 | $2 | $4 | $10 |
| Output price, $ per million tokens | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0% (high effort) | 65.2% | 72.7% | 70.0% |
| EEBench (electrical engineering) | 64.0% | 53.0% | 39.4% | 56.4% |
| AA-Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 37.6% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
Source: SpaceXAI launch page, accessed September 23, 2026.[1]
Grok 4.7 has the top score in two of the seven rows, EEBench and Harvey's Legal Agent Benchmark. Claude Fable 5.1 leads four, and GPT-5.6 Sol leads DeepSWE. The DeepSWE figure for Grok 4.7 is marked in the table as a high-effort score rather than xhigh. Against GPT-5.6 Sol, Grok 4.7 is ahead on five rows and behind on DeepSWE and HealthBench Professional.[1]
SpaceXAI changed one number after launch. At launch, both the benchmark image in its post on X and the table on the announcement page gave Grok 4.7 38.0% on Terminal-Bench 4.0. Archived copies of the page show 38.0% through at least 08:35 UTC on September 22. By September 23 the page said 37.6%, with no note of the change. All other cells are unchanged.[9][1] Coverage published on launch day, including The New Stack's, used the original 38.0% figure.[30]
The table's comparison set is also narrower than the rest of the post. It uses GPT-5.6 Sol rather than OpenAI's newer GPT-6 Astra, although a second chart on the same page includes Astra. In that chart, which covers three knowledge-work evaluations, Grok 4.7 is second each time: behind Fable 5.1 on GDPval (1,695 Elo against 1,735) and AA-Briefcase (1,657 against 1,678), and behind GPT-6 Astra on EEBench (64.0% against 69.3%). GPT-6 Astra scores 1,542 on GDPval and 1,569 on AA-Briefcase in the same chart.[1]
CursorBench cost chart
The announcement says that on CursorBench 4.0, "which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance." Its chart plots each model's score against the average cost per task at each reasoning setting.[1][10]
| Model | Low | Medium | High | Extra high | Max |
|---|---|---|---|---|---|
| Grok 4.7 | 33.1% ($1.58) | 41.6% ($3.49) | 43.9% ($4.69) | 46.3% ($6.01) | not offered |
| Claude Fable 5.1 | 45.1% ($5.44) | 46.8% ($7.05) | 49.2% ($9.08) | 51.6% ($13.01) | 51.8% ($17.28) |
| Claude Opus 5 | 40.7% ($4.87) | 43.3% ($6.94) | 44.7% ($9.00) | 46.1% ($11.43) | 46.6% ($11.95) |
| GPT-5.6 Sol | 24.6% ($0.87) | 31.1% ($1.77) | 35.7% ($2.85) | 37.7% ($4.40) | 41.7% ($8.23) |
| Claude Sonnet 5 | 24.1% ($1.39) | 28.0% ($2.31) | 30.8% ($3.48) | 32.0% ($4.55) | 34.1% ($7.17) |
Source: data labels of the CursorBench 4.0 chart on SpaceXAI's launch page.[1]
On these figures, Grok 4.7 scores higher than every other plotted configuration that costs less per task than it does, at each of its four settings. Claude Fable 5.1 still has the highest score overall. Its cheapest setting (45.1% at $5.44) falls between Grok 4.7's high and extra-high points.[1]
Safety claims
SpaceXAI says Grok 4.7 "was built with an entirely new safeguard stack" and is "the strongest model we've tested on refusals and jailbreak resistance." It says the model tops LatchBio's biosafety benchmark at 62.4%. It also says the model shows "the highest safety on HackerBench v0.3, our benchmark for risky and malicious cyber tasks, allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work." HackerBench is SpaceXAI's own benchmark. The company also said it had begun giving "select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research."[1] The launch post, model documentation and release notes do not link to a system card or separate safety report.[1][2][3]
Independent measurement
Artificial Analysis evaluated Grok 4.7 at xhigh effort on launch day and scored it 46 on Intelligence Index v4.3.2. That version combines ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.[16][18] The evaluator said the result brought SpaceXAI "into the top 4 AI labs." It is two points above Grok 4.6 (high) on the same index version, where Grok 4.6 scores 44. Scores on earlier index versions, such as the 61 that Grok 4.6 received at its launch on v4.1.1 (see the Grok 4.6 article), are not comparable.[16][20]
| Model (setting) | AA Intelligence Index v4.3.2 |
|---|---|
| Claude Opus 5.5 (max) | 57.6 |
| Claude Fable 5.1 (max) | 53.4 |
| GPT-6 Astra (max) | 52.7 |
| Claude Opus 5 (max) | 50.8 |
| Muse Spark 1.3 (max) | 48.1 |
| GPT-6 Sol (max) | 47.5 |
| GPT-5.6 Sol (max) | 47.0 |
| Grok 4.7 (xhigh) | 46.4 |
| Grok 4.7 (high) | 46.3 |
| MiMo-V2.6-Pro | 46.3 |
| Grok 4.6 (high) | 44.3 |
Source: Artificial Analysis model data, accessed September 23, 2026; scores rounded to one decimal place.[18][19][20]
Artificial Analysis ranks Grok 4.7 (xhigh) 21st of the 212 model configurations it compares it against. The high setting scores almost the same as xhigh.[18][19] In the evaluator's breakdown, the gains are concentrated in agentic knowledge work. Grok 4.7 gained 111 Elo over Grok 4.6 (high) on AA-Briefcase, reaching 1,657, and 90 Elo on GDPval-AA, reaching 1,695. "Outside of agentic knowledge work," Artificial Analysis wrote, it "broadly matches Grok 4.6 (high)." It improved by 4.5 percentage points on Terminal-Bench 4.0 and 3.0 points on GDP.pdf, and regressed by 3.7 points on AA-LCR and 1.1 points on AutomationBench-AA.[16][17] Artificial Analysis' own Terminal-Bench 4.0 run, in its standardized harness, recorded 25.8% for Grok 4.7 (xhigh). SpaceXAI's page now reports 37.6% on the same benchmark (38.0% at launch).[18][1]
Token use rose sharply. Artificial Analysis measured about 81,000 output tokens per Intelligence Index task for Grok 4.7 (xhigh), against 36,000 for Grok 4.6 (high) and 27,000 for GPT-6 Astra (max). That is 125% and 196% more, respectively.[16] Because the per-token prices did not change, the cost per task rose. The evaluator's model pages list $3.74 per Intelligence Index task for Grok 4.7 (xhigh) and $2.73 for Grok 4.7 (high), against $1.86 for Grok 4.6 (high). The totals are 240 million, 200 million and 94 million output tokens across the index.[18][19][20]
On speed, the same pages list a median output speed of 39 tokens per second for the xhigh configuration and 53 for high, against 59 for Grok 4.6 (high).[18][19][20] Artificial Analysis' benchmarking article gave a separate measure of about 188 tokens per second of answer output on long prompts, and an average of about 7.1 minutes per Intelligence Index task.[17] The evaluator also measured a lower hallucination rate on AA-Omniscience than for Grok 4.6 (high), 29% against 34%, with accuracy roughly unchanged.[17]
On the Artificial Analysis Coding Agent Index, which scores models inside their own coding agents rather than a standard harness, Grok 4.7 (xhigh) in Grok Build scored 56, up nine points from Grok 4.6 (xhigh). That placed it fourth among models in their native harnesses, behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5, and ahead of GPT-5.6 Sol. Its component scores rose from 65% to 73% on DeepSWE v1.1, from 18% to 33% on Terminal-Bench 4.0, and from 58% to 63% on SWE-Atlas-QnA.[16][17]
Other evaluations
The offensive-security company XBOW, which had early access and tested an early Grok 4.7 candidate, published results on launch day. In its production exploit-crafting harness (50 benchmarks, 10 repetitions each), Grok 4.7 was "slightly down from Grok 4.6 when both were given a fixed iteration budget." In broader agentic penetration tests, XBOW's systems built on Grok Build made 68 correct findings with Grok 4.7, against 42 with Grok 4.6. Its external harnesses made about 72 with Grok 4.6, and performance outside Build "moved slightly in the opposite direction." XBOW found xhigh reasoning "not distinguishable at all from high reasoning" in its tests. It also found that a failure mode in which Grok 4.6 got stuck reasoning without acting, seen in about 0.85% of runs, did not appear in any Grok 4.7 run it evaluated. It concluded that the model was "not quite GPT-6 or Mythos level quality, but it is getting closer."[31]
Reception
Coverage focused on the gap between the low token price and the benchmark position. The Decoder's headline said Grok 4.7 launched "at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6." The article noted that the $2 and $6 rates "are closer to Chinese models than Western frontier models" and cited the Artificial Analysis Terminal-Bench 4.0 results.[29] The New Stack's headline was "Grok 4.7 was built to work for hours. It still fails most of the time." The article described the harness-specific training as a possible source of lock-in: "Training models around specific tool schemas, context formats, and execution environments could make it harder for developers to swap models without sacrificing agent performance."[30] XBOW made a related observation from its own tests, writing that "the orchestration layer itself increasingly needs to fit the model."[31]
The launch results can be set against Musk's earlier claims. On August 12 he had said Grok 4.7 would "exceed all current models."[23] SpaceXAI's own table shows it leading on two of seven evaluations against GPT-5.6 Sol and Claude Fable 5.1. On the independent Artificial Analysis index, 20 configurations rank above it, among them Claude Fable 5.1, GPT-6 Astra, Claude Opus 5 and 5.5, Claude Fable 5, Muse Spark 1.3, OpenAI's GPT-6 Sol (released the day after Grok 4.7) and GPT-5.6 Sol.[18] SpaceXAI's case rests on price. The announcement's subtitle calls Grok 4.7 "Twice as fast, at half the price of comparable models," while the body says it is served "at the same price and speed as Grok 4.6."[1]
Successors: Grok 4.8, Grok 4.9 and Grok 5
Musk has described three later models on X. On September 14, 2026, a week before Grok 4.7 shipped, he wrote that "Grok 4.8, which is a 2.5T model trained with our new C++ software stack, will finish training this week and start RL."[27] Later that day he added: "Grok 4.8 will be a noticeable improvement. Grok 4.9 is probably Astra/Fable class. Grok 5 maybe better than anything. We shall see."[26] On the August 4 earnings call he had said Grok 5 "should be end of year" and would incorporate "the entire corpus of SpaceX data."[28] The C/C++ training stack and the rest of the roadmap are covered at Grok 4.6.
As of September 23, 2026, SpaceXAI's model and pricing documentation listed grok-4.7 as its newest text model, with no identifier, price or release note for Grok 4.8 or any later model. The Models page describes Grok 4.7 as "the most capable model we've built."[4] The parameter counts Musk has given, 2.1 trillion for Grok 4.7 and 2.5 trillion for Grok 4.8, do not appear in SpaceXAI's launch post or model documentation.[22][27][1][2]
See also
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20SpaceXAI. "Introducing Grok 4.7." September 21, 2026. Accessed September 23, 2026. x.ai/...grok-4-7 (archived copy with the original 38.0% Terminal-Bench figure: web.archive.org/...grok-4-7)
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8SpaceXAI. "Grok 4.7." SpaceXAI Docs. Accessed September 23, 2026. docs.x.ai/...grok-4-7
- ^1 ^2 ^3 ^4 ^5SpaceXAI. "Release Notes." SpaceXAI Docs. Accessed September 23, 2026. docs.x.ai/...release-notes
- ^1 ^2 ^3SpaceXAI. "Models." SpaceXAI Docs. Accessed September 23, 2026. docs.x.ai/...models
- ^1 ^2 ^3 ^4 ^5SpaceXAI. "Pricing." SpaceXAI Docs. Accessed September 23, 2026. docs.x.ai/...pricing
- ^1 ^2SpaceXAI. "Reasoning." SpaceXAI Docs. Accessed September 23, 2026. docs.x.ai/...reasoning
- ^1 ^2SpaceXAI. "Regional Endpoints." SpaceXAI Docs. Accessed September 23, 2026. docs.x.ai/...regions
- ^1 ^2SpaceXAI. "grok-4.7" model detail page. SpaceXAI Docs. Accessed September 23, 2026. docs.x.ai/...grok-4.7
- ^1 ^2 ^3SpaceXAI (@SpaceXAI). "Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed." X, September 21, 2026. x.com/...2102069815225586149
- ^1 ^2SpaceXAI (@SpaceXAI). "Grok 4.7 works longer on difficult tasks, checks its work more carefully, and comes with our strongest safeguards to date." X, September 21, 2026. x.com/...2102069817893150777
- ^SpaceXAI (@SpaceXAI). "Grok 4.7 is available now in Cursor and Grok Build, and the Grok API." X, September 21, 2026. x.com/...2102069822288720022
- ^Cursor Team. "Introducing Grok 4.7." Cursor, September 21, 2026. cursor.com/...grok-4-7
- ^GitHub. "Grok 4.7 is now available in GitHub Copilot." GitHub Changelog, September 21, 2026. github.blog/...-is-now-available-in-github-copilot
- ^OpenRouter. "SpaceXAI: Grok 4.7." Model listing, accessed September 23, 2026. openrouter.ai/...grok-4.7
- ^Elon Musk (@elonmusk). "Grok 4.7 is a strong combination of intelligence, speed & low cost." X, September 21, 2026. x.com/...2102071804495872374
- ^1 ^2 ^3 ^4 ^5 ^6Artificial Analysis (@ArtificialAnlys). "Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs." X, September 21, 2026. x.com/...2102074898327932987
- ^1 ^2 ^3 ^4 ^5Artificial Analysis. "Benchmarking Grok 4.7." September 21, 2026. artificialanalysis.ai/...benchmarking-grok-4-7
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Artificial Analysis. "Grok 4.7 (xhigh): Intelligence, Performance and Price Analysis." Accessed September 23, 2026. artificialanalysis.ai/...grok-4-7
- ^1 ^2 ^3 ^4Artificial Analysis. "Grok 4.7 (high): Intelligence, Performance and Price Analysis." Accessed September 23, 2026. artificialanalysis.ai/...grok-4-7-high
- ^1 ^2 ^3 ^4Artificial Analysis. "Grok 4.6 (high): Intelligence, Performance and Price Analysis." Accessed September 23, 2026. artificialanalysis.ai/...grok-4-6
- ^1 ^2 ^3Elon Musk (@elonmusk). "Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks." X, July 24, 2026. x.com/...2080724087593226311
- ^1 ^2 ^3 ^4Elon Musk (@elonmusk). "Grok 4.6 releases around August 7. This will be the 1.5T model with significantly improved SFT & RL. Grok 4.7 will be the 2.1T model released a few weeks later..." X, July 28, 2026. x.com/...2082123925283041545
- ^1 ^2 ^3 ^4Elon Musk (@elonmusk). "Grok 4.7 will exceed all current models..." Reply to @cognition, X, August 12, 2026. Shown by X as deleted by its author when checked on September 23, 2026. x.com/...2087606260539777263
- ^1 ^2 ^3Elon Musk (@elonmusk). "Grok 4.7 comes out in 10 days." X, September 2, 2026. x.com/...2094983639780204846
- ^1 ^2 ^3Elon Musk (@elonmusk). "Grok 4.7 needs a few more days to cook." Reply to @farzyness, X, September 11, 2026. x.com/...2098462085973741960
- ^1 ^2 ^3Elon Musk (@elonmusk). "Grok 4.7 should be roughly on par with Opus 5.0, not 5.1... Grok 5 maybe better than anything. We shall see." Reply to @itslueul, X, September 14, 2026. x.com/...2099458047408013751
- ^1 ^2Elon Musk (@elonmusk). "Grok 4.8, which is a 2.5T model trained with our new C++ software stack, will finish training this week and start RL." Reply to @techdevnotes, X, September 14, 2026. x.com/...2099308197802631191
- ^1 ^2Nehal Malik. "SpaceX Q2 Earnings Call Highlights: Starship, Starlink, AI & More." Not a Tesla App, August 5, 2026. notateslaapp.com/...tarship-starlink-grok-and-more
- ^1 ^2Matthias Bastian. "xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6." The Decoder, September 21, 2026. the-decoder.com/...-a-wide-gap-to-claude-and-gpt-6
- ^1 ^2 ^3Amanda Caswell. "Grok 4.7 was built to work for hours. It still fails most of the time." The New Stack, September 21, 2026. thenewstack.io/grok-4-7-agent-stamina
- ^1 ^2Albert Ziegler and Maria Knorps. "Grok 4.7 for Offensive Security: Orchestration Matters." XBOW, September 21, 2026. xbow.com/...grok-4-7-offensive-security-evaluation
- ^SpaceXAI (@SpaceXAI). "Compare Grok 4.7 (first) and 4.6 (second) building an open world city game." X, September 21, 2026. x.com/...2102069820418048091
- ^Elon Musk (@elonmusk). "SpaceX's massive corpus of world-class engineering data (excluding material blocked by ITAR) will be added during supplemental training of the 2T run." X, July 21, 2026. x.com/...2079446276299465185
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
2 revisions · v3 · 4,412 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: xg06 V2 independent verification (44 sources); 1 material (silent page edit 38.0->37.6) + 7 minor fixed 2026-09-23
Cite this page: AI Wiki. "Grok 4.7." aiwiki.ai, updated 23 Sept 2026, fact-checked 23 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/grok_4_7