Extended NYT Connections Benchmark
The Extended NYT Connections Benchmark evaluates large language models on 940 Connections puzzles with up to four added decoy words per puzzle. It is unaffiliated with the New York Times.[1]
Task
The supplied prompt asks for four distinct groups of four words sharing a concept. Connections can be literal categories or wordplay, including homophones, palindromes, and letter changes. The model must select exactly 16 words, even when the input contains more. It returns four comma-separated lines without category names or explanations.[2]
In the standard game, words can have several plausible associations. Finding a locally convincing group is not necessarily enough: the remaining words must also fit the intended solution. Research on Connections describes this combination of linguistic knowledge, context, and distractors as a test of abstract or lateral reasoning.[5]
Scoring
With g exact groups, the score is (g / 4)^2: 0%, 6.25%, 25%, 56.25%, or 100% for zero through four groups. The leaderboard averages puzzle scores, not complete-solve rates.[1]
The extended test permits one answer. A separate human-style simulation permits four mistakes and feedback, not this leaderboard's protocol.[1]
The public eval.py helper lowercases words and matches equal-length word sets against the answer groups. It strips formatting from responses and awards one quarter per matched group.[3]
That linear calculation does not implement the documented current quadratic headline score.[1][3]
October 2026 results
The October 6, 2026 update added three runs; all comparisons below cover 940 puzzles.[1]
Maintainer Lech Mazur highlighted these comparisons in his October 8 announcement:[6]
| Model and listed setting | Score | Earlier comparison | Score |
|---|---|---|---|
| GPT-6.1 Sol, high reasoning | 95.5% | GPT-6 Sol, high reasoning | 90.1% |
| Claude Sonnet 5.5, high reasoning | 80.5% | Claude Sonnet 5, high reasoning | 75.1% |
| Mistral Large 4, high | 27.4% | Mistral Large 3, non-reasoning | 7.5% |
The announcement reported about 54% lower per-puzzle cost for GPT-6.1 Sol than GPT-6 Sol, and about 87% lower cost for Sonnet 5.5 than Sonnet 5. It described Mistral Large 4's cost as substantially higher than its non-reasoning predecessor. These are the author's reported comparisons, not an independent reproduction.[6]
Cost comparisons use token usage and registered list/display prices, not invoices. Changing prices are averaged over a full pricing cycle.[1]
Research context and limits
This measures verbal puzzle solving, not general intelligence or software-engineering ability.[1]
Separate research should not be treated as an earlier version of this leaderboard. Samdarshi and colleagues' EMNLP 2024 study used 438 Connections games and compared models with novice and expert people. Their knowledge taxonomy distinguished ordinary semantic relations from encyclopedic knowledge, multiword expressions, and combinations of word form and meaning. That study found different difficulty patterns across those types.[4]
Todd and colleagues' Missed Connections study examined both interactive play and a simultaneous-answer variant. They found that presenting words already grouped by the solution could inflate results: a replication changed from 3.33% to 65% success when the word order supplied the groupings. This illustrates why presentation and protocol matter when comparing Connections evaluations; it is not an allegation about this benchmark.[5]
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Lech Mazur. Extended Connections README. Latest listed update October 6, 2026; accessed October 10, 2026.
- ^Lech Mazur. Extended puzzle prompt. Accessed October 10, 2026.
- ^1 ^2Lech Mazur. Public evaluation helper. Accessed October 10, 2026.
- ^Prisha Samdarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf, Tuhin Chakrabarty, and Smaranda Muresan. Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game. EMNLP, November 2024. DOI: 10.18653/v1/2024.emnlp-main.1182.
- ^1 ^2Graham Todd, Tim Merino, Sam Earle, and Julian Togelius. Missed Connections: Lateral Thinking Puzzles for Large Language Models. arXiv:2404.11730v2, April 21, 2024.
- ^1 ^2Lech Mazur. October 8 model and per-puzzle cost comparison. October 8, 2026.
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v1 · 623 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent full-article review against 6 cited primary and academic sources, October 10, 2026. Checked subject identity, specifications, availability, benchmark conditions and limitations.
Cite this page: AI Wiki. "Extended NYT Connections Benchmark." aiwiki.ai, updated 10 Oct 2026, fact-checked 10 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/extended_nyt_connections_benchmark