AI gold-medal result at the 2026 IOI
In September 2026, researchers at NVIDIA reported that a fine-tuned version of Nemotron 3 Ultra had scored 535.4 out of 600 points on the problem set of the 2026 International Olympiad in Informatics (IOI)
Explore Reasoning Models through related topics and the articles other pages reference most.
Ranked by links from other AI Wiki pages.
Articles that also belong to these categories. Counts cover all of Reasoning Models.
Showing 1-60 of 70 articles
In September 2026, researchers at NVIDIA reported that a fine-tuned version of Nemotron 3 Ultra had scored 535.4 out of 600 points on the problem set of the 2026 International Olympiad in Informatics (IOI)
ARC-AGI 1, short for Abstraction and Reasoning Corpus for Artificial General Intelligence, version 1
Adaptive thinking is an inference-time reasoning mode in the Anthropic Messages API in which a claude model decides, on a per-request basis, whether to use extended thinking at all and how much of it to spend
Agent planning refers to the process by which an AI agent determines a sequence of actions to accomplish a goal.
AlphaGeometry is a neuro-symbolic artificial intelligence system developed by Google DeepMind that solves competition-level geometry problems at a standard comparable to human Olympiad gold medalists.
AlphaProof is a reinforcement-learning system from Google DeepMind that finds and verifies formal mathematical proofs in the Lean 4 theorem prover
BIG-Bench Extra Hard (BBEH) is a reasoning benchmark released by Google DeepMind in February 2025 that replaces each of the 23 tasks in BIG-Bench Hard (BBH) with a new
As of July 2026, the strongest general reasoning models are Anthropic's Claude Opus 4.8 and Claude Fable 5, OpenAI's GPT-5.5, and Google's Gemini 3.1 Pro and Gemini 3 Pro in its Deep Think mode
Claude Opus 5 is a large language model released by Anthropic on 24 July 2026, the newest entry in the Claude Opus line and the successor to Claude Opus 4.8.
Command A Reasoning is an enterprise reasoning model released by Cohere on August 21, 2025, as part of the Command A family of large language models.
Commonsense reasoning is the ability to make the everyday, mostly tacit assumptions that ordinary humans take for granted, the implicit knowledge about how the physical world behaves, how minds work, how time…
DeepSeek V3.1 is a large language model developed by DeepSeek, released on August 19, 2025 and made broadly available via the official API on August 21, 2025.
DeepSeek-Prover is a family of open-weight large language models developed by Chinese AI laboratory DeepSeek for formal theorem proving in the Lean 4 proof assistant.
DeepSeek-R1 is an open-weight reasoning model and large language model family developed by the Chinese artificial intelligence laboratory DeepSeek. The original model was released on January 20, 2025.
DeepSeek-R1-Distill is a family of six open-weight reasoning language models released by DeepSeek on January 20, 2025, alongside the flagship DeepSeek-R1 reasoning model.
DeepSeekMath is a family of open-weight large language models specialized for mathematical reasoning, released by Chinese AI laboratory DeepSeek in February 2024.
Extended thinking is the product name Anthropic gives to the reasoning mode in its Claude family of large language models
GPQA (Graduate-Level Google-Proof Q&A) is a benchmark of expert-written, four-option multiple-choice questions in biology, physics, and chemistry.
GPT-5 Pro is a large language model developed by OpenAI and the highest-capability variant of the GPT-5 model family.
GPT-5.6 is a family of proprietary multimodal large language models developed by OpenAI. The family entered a limited preview on June 26, 2026, and became generally available on July 9, 2026.
GPT-6 Astra is an OpenAI model that the company began rolling out on September 3, 2026, describing it as "the world's most intelligent and aligned model" and as the successor to the GPT-5.6 family.
Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm for fine-tuning large language models that eliminates the separate critic (value) network used by PPO
GSM8K (Grade School Math 8K) is an English-language benchmark of grade-school arithmetic word problems released by OpenAI researchers in 2021.
Gemini 2.0 Flash Thinking is an experimental reasoning model released by Google as part of the Gemini 2.0 family.
Gemini 2.5 Deep Think is Google DeepMind's enhanced reasoning mode for the Gemini 2.5 Pro model that uses a technique called "parallel thinking" to explore many candidate solution paths at once before…
Goedel-Prover is an open-source large language model designed for automated formal theorem proving in Lean 4.
Grok 4 is a large language model developed by xAI and released on July 9, 2025. It is the fourth major generation of the Grok model family and was positioned as xAI's most capable model to date at its release.
Grok 4.5 is a proprietary multimodal large language model and reasoning model in the Grok family. It was developed by SpaceXAI in collaboration with Cursor and released through the xAI API on July 8, 2026.
Grok 4.6 is a proprietary large language model and reasoning model in the Grok family, developed by SpaceXAI and released jointly with Cursor through the xAI API on August 12, 2026.
Inference-time scaling (also called test-time compute scaling) is the practice of improving an AI model's output quality by allocating more computational resources during inference rather than during training.
Interleaved thinking is a feature of the Anthropic Messages API that lets Claude produce extended reasoning blocks between tool calls inside a single multi-turn agentic loop
Kimi K2 Thinking is a reasoning and agentic large language model released by the Chinese startup Moonshot AI on November 6, 2025.
Llama Nemotron is a family of open reasoning large language models built by Nvidia by post-training Meta's Llama models for math, coding, and agentic tasks.
MAI-Thinking-1 is a reasoning model developed by Microsoft AI, unveiled at Microsoft Build 2026 on June 2, 2026 as the company's first in-house flagship reasoning system.
MATH is a benchmark of 12,500 competition mathematics problems used to evaluate the mathematical problem-solving ability of machine learning systems, particularly large language models.
Magistral is the first family of reasoning models from Mistral AI, the French AI company, first released on June 10, 2025.
Marco-o1 is an open reasoning model released in November 2024 by the MarcoPolo team at Alibaba International Digital Commerce (AIDC).
MathArena is a public, continuously updated leaderboard and evaluation platform that measures the performance of large language models on mathematics competition problems released after each model's training…
Mathematical reasoning in AI is the ability of computer systems to solve mathematical problems: carrying out multi-step calculations, proving theorems, and answering competition or research questions that…
MiniMax M1 (stylised MiniMax-M1) is an open-weight large language reasoning model released on 16 June 2025 by the Shanghai-based artificial-intelligence company MiniMax
MiniMax M2.7 is a large language model released by the Chinese AI company MiniMax on March 18, 2026.
MuSR (Multistep Soft Reasoning) is a benchmark for evaluating multistep reasoning in large language models, built around long free-text narratives such as murder mysteries, object-placement scenarios, and…
Muse Spark is a proprietary multimodal reasoning model developed by Meta Superintelligence Labs (MSL), the artificial intelligence division Meta reorganized in 2025.
The NVIDIA Nemotron Model Reasoning Challenge was a Kaggle competition run by NVIDIA from March to June 2026 in which participants tried to improve the reasoning accuracy of a fixed open model, Nemotron 3 Nano…
Natural language inference (NLI), also known as recognising textual entailment (RTE), is the natural language processing task of deciding whether a hypothesis sentence is entailed by, contradicts, or is…
OLMo 3 is the third generation of fully open language models released by the Allen Institute for AI (Ai2).
The OpenAI o-series is a family of large language models developed by OpenAI that are trained with reinforcement learning to reason through an internal chain-of-thought before answering, making them OpenAI's…
OpenAI o1 is a family of proprietary large language models developed by OpenAI and trained to use additional computation before returning an answer.
OpenAI o1-mini is a smaller, faster, and cheaper reasoning model released by OpenAI on September 12, 2024, alongside o1-preview, and optimized for science, technology, engineering, and mathematics (STEM) tasks…
OpenAI o1-pro is the highest-compute variant of OpenAI's o1 reasoning model, designed to spend more inference-time compute so it "thinks harder" and returns the most reliable answers on the hardest…
OpenAI o3 is a family of reasoning-focused large language models developed by OpenAI and the second generation of the company's o-series reasoning models, best known for scoring 87.5% on the ARC-AGI…
OpenAI o3-mini is a reasoning-focused large language model released by OpenAI on January 31, 2025, the second commercial member of the o-series after OpenAI o1 and a smaller, cheaper
OpenAI o3-pro is a high-compute reasoning large language model released by OpenAI on June 10, 2025, designed as the professional, higher-reliability variant of the company's o3 reasoning model.
Phi-4-reasoning is a 14 billion parameter open weight reasoning model released by Microsoft Research on April 30, 2025.
Phi-4-mini-flash-reasoning is a 3.8 billion parameter open weight reasoning model released by Microsoft in July 2025.
ProcessBench is a benchmark for step-level verification of mathematical reasoning, built by the Qwen Team at Alibaba and released in December 2024.
QvQ (styled QVQ) is a family of experimental visual reasoning models from the Qwen team at Alibaba.
QwQ is a family of open-weight reasoning models from the Qwen team at Alibaba Cloud, built to compete with OpenAI's o1 and DeepSeek-R1 at a fraction of their size.
Qwen3.8 is the name Qwen uses for a model generation that includes hosted services and downloadable checkpoints.
Reinforcement Learning with Verifiable Rewards (RLVR) is a post-training paradigm for large language models in which the reward signal comes from a deterministic