# ChatGPT vs Claude vs Gemini vs Grok

> Source: https://aiwiki.ai/wiki/chatgpt_vs_claude_vs_gemini_vs_grok
> Updated: 2026-08-01
> Fact-checked: 2026-08-01
> Categories: AI Models, Conversational AI, Large Language Models
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "ChatGPT vs Claude vs Gemini vs Grok." aiwiki.ai, 1 Aug 2026. https://aiwiki.ai/wiki/chatgpt_vs_claude_vs_gemini_vs_grok
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**Comparison date:** July 28, 2026. Product access, model routing, prices, and features can change. This article separates the consumer products from the models and APIs available through them.

**[ChatGPT](https://aiwiki.ai/wiki/chatgpt)**, **[Claude](https://aiwiki.ai/wiki/claude)**, **[Gemini](https://aiwiki.ai/wiki/gemini)**, and **[Grok](https://aiwiki.ai/wiki/grok)** are generative AI assistants developed by [OpenAI](https://aiwiki.ai/wiki/openai), [Anthropic](https://aiwiki.ai/wiki/anthropic), [Google](https://aiwiki.ai/wiki/google), and [SpaceXAI](https://aiwiki.ai/wiki/xai), respectively. Each product combines one or more language [models](https://aiwiki.ai/wiki/model) with a user interface, tools, usage limits, and integrations. A comparison of the products is therefore not the same as a comparison of four fixed models.

There is no supportable universal ranking of the four assistants. Results depend on the exact model and reasoning setting, the prompt, available tools, account plan, region, safety routing, and evaluation date. A product can also use different models for a quick chat, an extended reasoning request, a coding task, or a research workflow. The most defensible way to choose among them is to identify the required work surface and then test representative tasks under controlled conditions.

## Product snapshot on July 28, 2026

The following table records provider-documented availability by the comparison date. It does not treat a model announced for an API or coding product as automatically available in every consumer chat.

| Product | Provider-documented model state by the cutoff | Availability stated by the provider |
| --- | --- | --- |
| ChatGPT | OpenAI made the GPT-5.6 family generally available on July 9. In ChatGPT, Plus, Pro, Business, and Enterprise users could access GPT-5.6 Sol through medium and higher effort settings; Pro and Enterprise users could also select Sol Pro. | OpenAI said GPT-5.6 was available across ChatGPT, Codex, and the OpenAI API, with a global rollout beginning on launch day.[1] |
| Claude | Anthropic released Claude Opus 5 on July 24. It described Opus 5 as the new default on Claude Max and the strongest model on Claude Pro. | Anthropic said Opus 5 was available on all Claude platforms and through the Claude API.[2] |
| Gemini | Google released Gemini 3.6 Flash on July 21 and said Gemini 3.5 Pro was still testing with partners. Gemini 3.1 Pro, released in February, remained Google's documented Pro model. | Google made 3.6 Flash available in the Gemini app, Gemini API, AI Studio, Android Studio, Antigravity, and enterprise products.[3] Google's 3.1 Pro model card listed the Gemini app, API, AI Studio, Antigravity, Vertex AI, Gemini Enterprise, and NotebookLM as distribution channels.[4] |
| Grok | SpaceXAI released Grok 4.5 on July 16. Its launch notice described it as the default in Grok Build. | The notice documented availability in Grok Build, Cursor, and the SpaceXAI API. It did not say that Grok 4.5 was the default model in the ordinary Grok chat product, so that should not be inferred from the launch notice.[5] |

This snapshot also shows why model names must accompany product comparisons. A test labeled only "ChatGPT versus Claude" does not disclose whether it compared GPT-5.6 Sol with Opus 5, another model selected by a router, or different reasoning settings. The same problem applies when a consumer plan has rate limits or automatically falls back to another model.

## Research, retrieval, and connected data

All four product families had provider-documented ways to work with information beyond a model's training data, but their research modes and connected sources were not identical.

OpenAI's deep research feature performs multi-step web research and returns a report with source links. By February 2026, OpenAI also said the feature could connect to apps or Model Context Protocol services and restrict web searches to selected sites.[6] Anthropic's Research feature could search the web and connected work sources, while its Integrations system added remote services through the Model Context Protocol. Anthropic said Research reports included citations and that access had expanded to Pro, Max, Team, and Enterprise plans by June 2025.[7]

Google introduced Deep Research in Gemini as a workflow that explores a topic and compiles a report with links to original sources.[8] SpaceXAI introduced Grok connectors on web, iOS, and Android in May 2026, including integrations for SharePoint, Outlook, OneDrive, Google Workspace, Notion, GitHub, and Linear, plus custom Model Context Protocol servers.[9]

These features can improve access to current or private information, but retrieval is not a guarantee of factual accuracy. A research answer can still select a weak source, misread a document, omit contrary evidence, or attach a citation that does not support the nearby sentence. For consequential work, the cited source should be opened and checked directly.

## Coding and task-execution surfaces

The four providers also offered different surfaces for software and multi-step work by the cutoff:

- OpenAI made GPT-5.6 available in Codex as well as ChatGPT and the API. Its launch announcement described model choices and effort controls in Codex.[1]
- Anthropic made Opus 5 available across its platforms, including Claude Code, and documented fallback behavior for some requests affected by safeguards.[2]
- Google distributed Gemini models through Antigravity, AI Studio, Android Studio, and the Gemini API, in addition to the Gemini app.[3][4]
- SpaceXAI positioned Grok 4.5 as the default model in Grok Build and offered it through the API and Cursor.[5]

These products should be tested in the surface that will actually be used. A chat response that produces a code block is not equivalent to an agent working in a repository with permission to inspect files, run tests, or edit code. Similarly, a demonstration that uses private tools does not establish the same result in a plain chat session.

## Multimodal input and files

File and media support is another product-level distinction. For example, Google's February 2026 model card documented text, image, audio, and video inputs for Gemini 3.1 Pro, a context window of up to one million tokens, and text output up to 64,000 tokens.[4] Those specifications describe that model, not every Gemini mode or account limit.

Comparable care is needed for the other assistants. The model's accepted input types, the product's upload interface, per-file limits, extraction software, and available output formats are separate layers. A model that can interpret an image through an API may not have the same image, audio, or document workflow in every consumer plan. Users comparing document analysis should test their own file types, page counts, charts, tables, and scanned material rather than infer performance from a model-family label.

## API price snapshot

The figures below are uncached list prices per one million input and output tokens stated in provider launch announcements by the cutoff. They are included for dated reference, not as a price-performance ranking. The rows are not matched capability tiers, and they exclude subscription fees, tool charges, caching rules, long-context premiums, batch discounts, and the number of tokens required to finish a task.

| Model | Announcement date | Input | Output | Important scope |
| --- | --- | ---: | ---: | --- |
| GPT-5.6 Sol | July 9, 2026 | US$5 | US$30 | OpenAI's flagship Sol API model at launch.[1] |
| Claude Opus 5 | July 24, 2026 | US$5 | US$25 | Anthropic's standard Opus 5 API rate; Fast mode was listed at twice the base rate.[2] |
| Gemini 3.6 Flash | July 21, 2026 | US$1.50 | US$7.50 | Google's Flash model, designed as an efficient workhorse rather than a Pro-tier match.[3] |
| Grok 4.5 | July 16, 2026 | US$2 | US$6 | The list price in SpaceXAI's Grok 4.5 launch notice.[5] |

Consumer subscriptions cannot be compared by token price alone. They bundle different model limits, tools, storage, integrations, and priority levels. Published monthly prices can also vary by country, tax treatment, billing channel, and promotion. A current purchase decision should therefore use the checkout page for the buyer's region rather than this historical snapshot.

## Why benchmark tables do not identify one winner

Benchmark results can answer a narrow question when the task, dataset, model version, prompt, sampling method, tools, and scoring procedure are disclosed. They do not establish that one assistant is better for every use.

The Holistic Evaluation of Language Models project was created in part because language models had been evaluated on sparse and inconsistent sets of scenarios. Its authors used multiple scenarios and metrics, including accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency, to expose trade-offs that a single score can hide.[10] That approach supports evaluating the properties relevant to a real use case rather than collapsing them into an undefined "best" label.

Human-preference leaderboards measure something different from factual correctness. Chatbot Arena uses pairwise votes from users to estimate preferences between model responses.[11] Those votes can be useful for conversational quality, but a preferred answer can still contain a factual or citation error. Automated judging also has known limits. Research on MT-Bench and language-model judges found position, verbosity, and self-enhancement biases, even while reporting substantial agreement with human preferences.[12]

Vendor benchmark charts need additional caution. Providers may use different harnesses, reasoning budgets, tools, fallback rules, or estimates of competitor cost. A result from one provider should be presented as that provider's evaluation unless an independent evaluator reproduced it under a common protocol. For this reason, this article does not carry forward the earlier page's estimated hallucination percentages or universal rankings.

## A controlled way to compare the assistants

A useful evaluation can be small, but it should be reproducible.

1. **Define the tasks.** Use examples from the intended workload, such as source-grounded research, data extraction, coding, long-document review, image interpretation, or routine writing.
2. **Freeze the setup.** Record the date, region, account plan, product surface, exact model label, reasoning setting, enabled tools, connected data, system instructions, and prompt.
3. **Use answer keys where possible.** For extraction, calculation, coding, and factual research, define what counts as correct before viewing the outputs.
4. **Repeat variable tasks.** Run more than one trial when sampling or routing can change the answer. Rotate response order if people are judging outputs side by side.
5. **Score separate dimensions.** Accuracy, unsupported claims, citation support, instruction following, latency, editing effort, and total cost should not be merged unless the weighting is explicit.
6. **Test failure cases.** Include missing information, ambiguous instructions, adversarial documents, outdated premises, and requests that should produce uncertainty rather than a confident guess.
7. **Recheck over time.** Model aliases, routing, safeguards, and tools can change without the consumer product changing its name.

For sensitive personal, legal, medical, financial, or company data, the comparison should also cover the applicable contract and privacy settings, retention, training use, access controls, audit features, data location, and administrator policies. Consumer, business, enterprise, and API terms should not be assumed to be interchangeable.

## Selection considerations

The most suitable product is the one that meets the required workflow under the user's own tests and constraints. Relevant questions include:

- Does the needed model exist on the buyer's plan and in the buyer's region?
- Does the work happen mainly in a chat, a coding agent, an office suite, a mobile app, or an API?
- Which web, file, email, calendar, repository, and enterprise-data connections are required?
- Are citations needed, and do they consistently support the generated claims?
- What limits apply to files, context, reasoning, tool calls, and repeated use?
- What is the measured cost per completed task, including review and correction time?
- Do the privacy, security, and administrative controls meet the organization's requirements?

These questions lead to a defensible shortlist without claiming that ChatGPT, Claude, Gemini, or Grok is universally superior. Because the products change quickly, any recommendation should include its test date and exact configuration.

## See also

- [AI model comparison finder](https://aiwiki.ai/tools/compare_models)

## References

1. OpenAI. "GPT-5.6: Frontier intelligence that scales with your ambition." July 9, 2026. https://openai.com/index/gpt-5-6/

2. Anthropic. "Introducing Claude Opus 5." July 24, 2026. https://www.anthropic.com/news/claude-opus-5

3. Google. "Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber." July 21, 2026. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/

4. Google DeepMind. "Gemini 3.1 Pro" model card. Published February 19, 2026. https://deepmind.google/models/model-cards/gemini-3-1-pro

5. SpaceXAI. "Introducing Grok 4.5." July 16, 2026. https://x.ai/news/grok-4-5

6. OpenAI. "Introducing deep research." February 2, 2025, updated February 10, 2026. https://openai.com/index/introducing-deep-research/

7. Anthropic. "Claude can now connect to your world." May 1, 2025, availability update June 3, 2025. https://www.anthropic.com/news/integrations

8. Google. "Try Deep Research and our new experimental model in Gemini, your AI assistant." December 11, 2024. https://blog.google/products-and-platforms/products/gemini/google-gemini-deep-research/

9. SpaceXAI. "Connectors in web, iOS, and Android." May 6, 2026. https://x.ai/news/grok-connectors

10. Percy Liang et al. "Holistic Evaluation of Language Models." Transactions on Machine Learning Research, 2023. https://arxiv.org/abs/2211.09110

11. Wei-Lin Chiang et al. "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference." arXiv:2403.04132, 2024. https://arxiv.org/abs/2403.04132

12. Lianmin Zheng et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." NeurIPS 2023 Datasets and Benchmarks Track. https://arxiv.org/abs/2306.05685

