Best AI Models for Reasoning and Math
As of July 2026, the strongest general reasoning models are Anthropic's Claude Opus 4.8 and Claude Fable 5, OpenAI's GPT-5.5, and Google's Gemini 3.1 Pro and Gemini 3 Pro in its Deep Think mode
Explore Reasoning Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Reasoning Models.
Showing 1-39 of 39 articles
As of July 2026, the strongest general reasoning models are Anthropic's Claude Opus 4.8 and Claude Fable 5, OpenAI's GPT-5.5, and Google's Gemini 3.1 Pro and Gemini 3 Pro in its Deep Think mode
Claude Opus 5 is a large language model released by Anthropic on 24 July 2026, the newest entry in the Claude Opus line and the successor to Claude Opus 4.8.
Command A Reasoning is an enterprise reasoning model released by Cohere on August 21, 2025, as part of the Command A family of large language models.
DeepSeek V3.1 is a large language model developed by DeepSeek, released on August 19, 2025 and made broadly available via the official API on August 21, 2025.
DeepSeek-Prover is a family of open-weight large language models developed by Chinese AI laboratory DeepSeek for formal theorem proving in the Lean 4 proof assistant.
DeepSeek-R1 is an open-weight reasoning model and large language model family developed by the Chinese artificial intelligence laboratory DeepSeek. The original model was released on January 20, 2025.
DeepSeek-R1-Distill is a family of six open-weight reasoning language models released by DeepSeek on January 20, 2025, alongside the flagship DeepSeek-R1 reasoning model.
DeepSeekMath is a family of open-weight large language models specialized for mathematical reasoning, released by Chinese AI laboratory DeepSeek in February 2024.
GPT-5 Pro is a large language model developed by OpenAI and the highest-capability variant of the GPT-5 model family.
GPT-5.6 is a family of proprietary multimodal large language models developed by OpenAI. The family entered a limited preview on June 26, 2026, and became generally available on July 9, 2026.
GPT-6 Astra is an OpenAI model that the company began rolling out on September 3, 2026, describing it as "the world's most intelligent and aligned model" and as the successor to the GPT-5.6 family.
GSM8K (Grade School Math 8K) is an English-language benchmark of grade-school arithmetic word problems released by OpenAI researchers in 2021.
Gemini 2.0 Flash Thinking is an experimental reasoning model released by Google as part of the Gemini 2.0 family.
Gemini 2.5 Deep Think is Google DeepMind's enhanced reasoning mode for the Gemini 2.5 Pro model that uses a technique called "parallel thinking" to explore many candidate solution paths at once before…
Goedel-Prover is an open-source large language model designed for automated formal theorem proving in Lean 4.
Grok 4 is a large language model developed by xAI and released on July 9, 2025. It is the fourth major generation of the Grok model family and was positioned as xAI's most capable model to date at its release.
Grok 4.5 is a proprietary multimodal large language model and reasoning model in the Grok family. It was developed by SpaceXAI in collaboration with Cursor and released through the xAI API on July 8, 2026.
Grok 4.6 is a proprietary large language model and reasoning model in the Grok family, developed by SpaceXAI and released jointly with Cursor through the xAI API on August 12, 2026.
Kimi K2 Thinking is a reasoning and agentic large language model released by the Chinese startup Moonshot AI on November 6, 2025.
Llama Nemotron is a family of open reasoning large language models built by Nvidia by post-training Meta's Llama models for math, coding, and agentic tasks.
Magistral is the first family of reasoning models from Mistral AI, the French AI company, first released on June 10, 2025.
MiniMax M1 (stylised MiniMax-M1) is an open-weight large language reasoning model released on 16 June 2025 by the Shanghai-based artificial-intelligence company MiniMax
MiniMax M2.7 is a large language model released by the Chinese AI company MiniMax on March 18, 2026.
OLMo 3 is the third generation of fully open language models released by the Allen Institute for AI (Ai2).
The OpenAI o-series is a family of large language models developed by OpenAI that are trained with reinforcement learning to reason through an internal chain-of-thought before answering, making them OpenAI's…
OpenAI o1 is a family of proprietary large language models developed by OpenAI and trained to use additional computation before returning an answer.
OpenAI o1-mini is a smaller, faster, and cheaper reasoning model released by OpenAI on September 12, 2024, alongside o1-preview, and optimized for science, technology, engineering, and mathematics (STEM) tasks…
OpenAI o1-pro is the highest-compute variant of OpenAI's o1 reasoning model, designed to spend more inference-time compute so it "thinks harder" and returns the most reliable answers on the hardest…
OpenAI o3 is a family of reasoning-focused large language models developed by OpenAI and the second generation of the company's o-series reasoning models, best known for scoring 87.5% on the ARC-AGI…
OpenAI o3-mini is a reasoning-focused large language model released by OpenAI on January 31, 2025, the second commercial member of the o-series after OpenAI o1 and a smaller, cheaper
OpenAI o3-pro is a high-compute reasoning large language model released by OpenAI on June 10, 2025, designed as the professional, higher-reliability variant of the company's o3 reasoning model.
Phi-4-reasoning is a 14 billion parameter open weight reasoning model released by Microsoft Research on April 30, 2025.
Phi-4-mini-flash-reasoning is a 3.8 billion parameter open weight reasoning model released by Microsoft in July 2025.
QwQ is a family of open-weight reasoning models from the Qwen team at Alibaba Cloud, built to compete with OpenAI's o1 and DeepSeek-R1 at a fraction of their size.
Qwen3.8 is the name Qwen uses for a model generation that includes hosted services and downloadable checkpoints.
Self-consistency is a decoding strategy for large language models that samples multiple chain-of-thought reasoning paths for the same question and returns the answer that the majority of those paths agree on
Test-time compute (also called inference-time compute scaling or test-time scaling) is the practice of allocating additional computation while a large language model answers a query
ZAYA1-8B is an open-weight, reasoning-focused Mixture-of-Experts (MoE) large language model released by San Francisco-based AI research lab Zyphra on May 6, 2026.
OpenAI o4-mini is a compact reasoning model developed by OpenAI, released on April 16, 2025.