Reasoning Models

Explore Reasoning Models through related topics and the articles other pages reference most.

Most referenced in this topic

Ranked by links from other AI Wiki pages.

Explore articles

Browse subtopics (32)

Articles that also belong to these categories. Counts cover all of Reasoning Models.

Showing 1-60 of 70 articles

Adaptive thinking

Adaptive thinking is an inference-time reasoning mode in the Anthropic Messages API in which a claude model decides, on a per-request basis, whether to use extended thinking at all and how much of it to spend

AI InferenceAnthropic

AlphaGeometry

AlphaGeometry is a neuro-symbolic artificial intelligence system developed by Google DeepMind that solves competition-level geometry problems at a standard comparable to human Olympiad gold medalists.

AI ModelsAI for Science

AlphaProof

AlphaProof is a reinforcement-learning system from Google DeepMind that finds and verifies formal mathematical proofs in the Lean 4 theorem prover

AI ModelsAI for Science

Claude Opus 5

Claude Opus 5 is a large language model released by Anthropic on 24 July 2026, the newest entry in the Claude Opus line and the successor to Claude Opus 4.8.

AI ModelsAnthropic

Commonsense reasoning

Commonsense reasoning is the ability to make the everyday, mostly tacit assumptions that ordinary humans take for granted, the implicit knowledge about how the physical world behaves, how minds work, how time…

Artificial Intelligence

DeepSeek V3.1

DeepSeek V3.1 is a large language model developed by DeepSeek, released on August 19, 2025 and made broadly available via the official API on August 21, 2025.

AI ModelsChinese AI

DeepSeek-R1

DeepSeek-R1 is an open-weight reasoning model and large language model family developed by the Chinese artificial intelligence laboratory DeepSeek. The original model was released on January 20, 2025.

Chinese AILarge Language Models

DeepSeek-R1-Distill

DeepSeek-R1-Distill is a family of six open-weight reasoning language models released by DeepSeek on January 20, 2025, alongside the flagship DeepSeek-R1 reasoning model.

AI ModelsChinese AI

GPQA

GPQA (Graduate-Level Google-Proof Q&A) is a benchmark of expert-written, four-option multiple-choice questions in biology, physics, and chemistry.

AI Benchmarks

GPT-5.6

GPT-5.6 is a family of proprietary multimodal large language models developed by OpenAI. The family entered a limited preview on June 26, 2026, and became generally available on July 9, 2026.

AI ModelsLarge Language Models

GPT-6 Astra

GPT-6 Astra is an OpenAI model that the company began rolling out on September 3, 2026, describing it as "the world's most intelligent and aligned model" and as the successor to the GPT-5.6 family.

AI ModelsLarge Language Models

GRPO

Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm for fine-tuning large language models that eliminates the separate critic (value) network used by PPO

AI InferenceChinese AI

Grok 4

Grok 4 is a large language model developed by xAI and released on July 9, 2025. It is the fourth major generation of the Grok model family and was positioned as xAI's most capable model to date at its release.

AI CompaniesAI Models

Grok 4.5

Grok 4.5 is a proprietary multimodal large language model and reasoning model in the Grok family. It was developed by SpaceXAI in collaboration with Cursor and released through the xAI API on July 8, 2026.

AI ModelsGenerative AI

Grok 4.6

Grok 4.6 is a proprietary large language model and reasoning model in the Grok family, developed by SpaceXAI and released jointly with Cursor through the xAI API on August 12, 2026.

AI ModelsGenerative AI

Inference-time scaling

Inference-time scaling (also called test-time compute scaling) is the practice of improving an AI model's output quality by allocating more computational resources during inference rather than during training.

AI InferenceAI Research

Interleaved thinking

Interleaved thinking is a feature of the Anthropic Messages API that lets Claude produce extended reasoning blocks between tool calls inside a single multi-turn agentic loop

AI AgentsAnthropic

MAI-Thinking-1

MAI-Thinking-1 is a reasoning model developed by Microsoft AI, unveiled at Microsoft Build 2026 on June 2, 2026 as the company's first in-house flagship reasoning system.

Microsoft

MATH

MATH is a benchmark of 12,500 competition mathematics problems used to evaluate the mathematical problem-solving ability of machine learning systems, particularly large language models.

AI BenchmarksModel Evaluation

Marco-o1

Marco-o1 is an open reasoning model released in November 2024 by the MarcoPolo team at Alibaba International Digital Commerce (AIDC).

Chinese AIOpen Source AI

MathArena

MathArena is a public, continuously updated leaderboard and evaluation platform that measures the performance of large language models on mathematics competition problems released after each model's training…

AI BenchmarksArtificial Intelligence

MiniMax M1

MiniMax M1 (stylised MiniMax-M1) is an open-weight large language reasoning model released on 16 June 2025 by the Shanghai-based artificial-intelligence company MiniMax

Chinese AILarge Language Models

MuSR

MuSR (Multistep Soft Reasoning) is a benchmark for evaluating multistep reasoning in large language models, built around long free-text narratives such as murder mysteries, object-placement scenarios, and…

AI BenchmarksModel Evaluation

Muse Spark

Muse Spark is a proprietary multimodal reasoning model developed by Meta Superintelligence Labs (MSL), the artificial intelligence division Meta reorganized in 2025.

Meta AIMultimodal AI

OpenAI o1-mini

OpenAI o1-mini is a smaller, faster, and cheaper reasoning model released by OpenAI on September 12, 2024, alongside o1-preview, and optimized for science, technology, engineering, and mathematics (STEM) tasks…

Large Language ModelsOpenAI

OpenAI o1-pro

OpenAI o1-pro is the highest-compute variant of OpenAI's o1 reasoning model, designed to spend more inference-time compute so it "thinks harder" and returns the most reliable answers on the hardest…

Large Language ModelsOpenAI

OpenAI o3

OpenAI o3 is a family of reasoning-focused large language models developed by OpenAI and the second generation of the company's o-series reasoning models, best known for scoring 87.5% on the ARC-AGI…

AI ModelsLarge Language Models

OpenAI o3-mini

OpenAI o3-mini is a reasoning-focused large language model released by OpenAI on January 31, 2025, the second commercial member of the o-series after OpenAI o1 and a smaller, cheaper

Large Language ModelsOpenAI

OpenAI o3-pro

OpenAI o3-pro is a high-compute reasoning large language model released by OpenAI on June 10, 2025, designed as the professional, higher-reliability variant of the company's o3 reasoning model.

Large Language ModelsOpenAI

QwQ

QwQ is a family of open-weight reasoning models from the Qwen team at Alibaba Cloud, built to compete with OpenAI's o1 and DeepSeek-R1 at a fraction of their size.

Chinese AILarge Language Models

Qwen3.8

Qwen3.8 is the name Qwen uses for a model generation that includes hosted services and downloadable checkpoints.

AI ModelsChinese AI

RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) is a post-training paradigm for large language models in which the reward signal comes from a deterministic

AI InferenceReinforcement Learning