AI gold-medal result at the 2026 IOI
In September 2026, researchers at NVIDIA reported that a fine-tuned version of Nemotron 3 Ultra had scored 535.4 out of 600 points on the problem set of the 2026 International Olympiad in Informatics (IOI)
Explore Reasoning Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Reasoning Models.
Showing 1-13 of 13 articles
In September 2026, researchers at NVIDIA reported that a fine-tuned version of Nemotron 3 Ultra had scored 535.4 out of 600 points on the problem set of the 2026 International Olympiad in Informatics (IOI)
ARC-AGI 1, short for Abstraction and Reasoning Corpus for Artificial General Intelligence, version 1
BIG-Bench Extra Hard (BBEH) is a reasoning benchmark released by Google DeepMind in February 2025 that replaces each of the 23 tasks in BIG-Bench Hard (BBH) with a new
GPQA (Graduate-Level Google-Proof Q&A) is a benchmark of expert-written, four-option multiple-choice questions in biology, physics, and chemistry.
GSM8K (Grade School Math 8K) is an English-language benchmark of grade-school arithmetic word problems released by OpenAI researchers in 2021.
MATH is a benchmark of 12,500 competition mathematics problems used to evaluate the mathematical problem-solving ability of machine learning systems, particularly large language models.
MathArena is a public, continuously updated leaderboard and evaluation platform that measures the performance of large language models on mathematics competition problems released after each model's training…
Mathematical reasoning in AI is the ability of computer systems to solve mathematical problems: carrying out multi-step calculations, proving theorems, and answering competition or research questions that…
MuSR (Multistep Soft Reasoning) is a benchmark for evaluating multistep reasoning in large language models, built around long free-text narratives such as murder mysteries, object-placement scenarios, and…
The NVIDIA Nemotron Model Reasoning Challenge was a Kaggle competition run by NVIDIA from March to June 2026 in which participants tried to improve the reasoning accuracy of a fixed open model, Nemotron 3 Nano…
Natural language inference (NLI), also known as recognising textual entailment (RTE), is the natural language processing task of deciding whether a hypothesis sentence is entailed by, contradicts, or is…
ProcessBench is a benchmark for step-level verification of mathematical reasoning, built by the Qwen Team at Alibaba and released in December 2024.
SimpleBench is a text-only benchmark for large language models created by Philip, the host of the AI Explained YouTube channel, with collaborator Hemang.