GDP.pdf
GDP.pdf is a benchmark from Surge AI that tests whether multimodal language models can answer realistic professional questions about real PDF documents, such as benefits packets, leases, datasheets, clinical guidelines, insurance policies and construction plans. It has 100 tasks, each pairing one PDF with a question written by a working professional. Each task comes with a rubric of atomic yes/no criteria, and the headline score is a strict pass rate: a model gets credit for a task only when it satisfies every criterion.[1][2] Surge released the benchmark on April 14, 2026, and describes it on its benchmark page as "a multimodal and reasoning benchmark that takes real-world prompts and PDFs pulled directly from expert professional workflows."[2][4] The accompanying paper, "GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents" (arXiv:2607.11192), is by Suhaas Garre, Emily Ritchie, Sushant Mehta and Edwin Chen of Surge AI.[1][2]
Scores on the benchmark are low. In Surge's April 2026 pilot, no model scored above 15%. As of September 30, 2026, the top entry on Surge's own leaderboard was GPT-6 Astra (max reasoning) at 34.2%.[3][4][6] GDP.pdf has also become a figure in lab marketing. OpenAI listed it in its GPT-5.6 and GPT-6.1 Sol launch posts, Anthropic reported it in four system cards between June and September 2026, Google put it in the Gemini 3.7 Flash and Gemini 3.8 Flash model cards, and Artificial Analysis added its own implementation to the Artificial Analysis Intelligence Index in September 2026.[14][15][17][18][19][20][23][25][30] These parties do not measure the benchmark the same way. They differ in harness, judge model, input handling, reasoning setting and metric, so the same model can have scores more than nine points apart depending on who ran it (see the Scores section below).
Background and name
The paper argues that document-AI benchmarks usually test one skill at a time: OCR, layout analysis, chart reasoning, table QA or document visual question answering. It says a high score on any of these "does not necessarily reveal whether a model can answer a realistic question that someone in the field would actually ask about a specific PDF."[1] The authors single out three problems: document structure (multi-page tables, sidebars, legends, footnotes, amendments appended at the end), background knowledge the document assumes, and failures that do not look like failures, where "the model cites a clause that exists and a number that is on the page; the clause is simply not the one that governs the user's question."[2]
The first arXiv version gives the reasoning behind the name: it refers to "the professional documents behind everyday economic activity ("GDP"), in the format they actually circulate in (".pdf")."[3] Surge's launch post starts with an anecdote. An emergency physician who works with Surge asked a frontier model for the correct treatment of an acute pulmonary embolism, which meant reading a PDF of 2026 treatment guidelines, and in Surge's telling "Every single frontier model failed."[5] The same post frames the stakes this way: "Before we trust enterprise agents with the workflows that run the economy, they have to master the paperwork that sustains it."[5]
The arXiv comments field says the paper was accepted at the 2nd Workshop on Knowledge-Intensive Multimodal Reasoning (KnowledgeMR) at CVPR 2026 under an earlier title, "PDFParse: A Benchmark for Grounded Multimodal Reasoning over Professional PDF Documents." The workshop is non-archival.[1] Version 1 was submitted on July 13, 2026. Versions 2 and 3 followed on July 14 and 15, and version 2 replaced the April pilot results with the July 2026 leaderboard of 17 models.[1][3]
Design
Task format
The paper defines each item as a PDF, a natural-language question, an expert rubric of atomic criteria, an expert reference answer, a domain label and a set of capability tags. The grader marks each rubric criterion as pass (1) or fail (0). An item's rubric score is the fraction of criteria it passes, and it passes strictly only if every criterion passes. The paper ranks models by strict pass rate and uses the mean rubric score for finer-grained analysis.[2] The reference answer helps the expert write and sanity-check the rubric, but the judge never sees it.[2]
The paper lists five design principles:[2]
| Principle | What the paper says |
|---|---|
| Prior knowledge | A task was rejected if general knowledge was enough to answer it; the answer must depend on the attached PDF |
| Realism | Questions came from contributors' own jobs, phrased as they would type them to an assistant mid-task |
| Adversarial construction | A candidate was kept only if at least two of several frontier models made a "major, meaningful error" on it |
| Knowledge-intensive grounding | Tasks where perception and domain knowledge interact (exclusions, footnotes, legends, plan symbols, amendments) were favored |
| Diagnostic granularity | Every item carries capability tags so results can be sliced by failure type |
Domains and documents
The 100 items are spread evenly across ten domains: Finance, Healthcare, Legal, STEM/Research, Engineering, Construction, Manufacturing/Supply Chain, Insurance, Real Estate and Human Resources.[2] The paper's coverage table pairs each domain with typical documents and typical failures:[2]
| Domain | Representative PDFs | How models typically fail (per the paper) |
|---|---|---|
| Finance | Earnings releases, investor filings, analyst materials | Values taken from the adjacent column or wrong note field; tables that cross pages lose alignment |
| Healthcare | Reviews, dosage tables, clinical guidelines | Losing place in a long review, skipping a figure footnote, answering when the information is absent |
| Legal | Contracts, leases, policy language, filings | Citing a real clause that is not the governing one; dropping cross-references |
| STEM/Research | Climate reports, technical reports, scientific figures | Misread legends and chart values; evidence from several figures never combined |
| Engineering | Datasheets and specification sheets | Log-scale plots misread; units mixed up; merged headers scramble lookups |
| Construction | Floor plans and schedules | Symbols not matched to the legend; schedule never reconciled with the drawing |
| Manufacturing | Process notes, packing lists, certificates of conformance | Footnotes skipped; model priors override the instructions |
| Insurance | Auto policies and endorsements | Reading stops at the insuring agreement, before the exclusions |
| Real Estate | Valuation reports, inspection reports, deeds, amendments | Status columns misread; superseded sections quoted as if in force |
| HR | Benefits packets, leave-policy tables, handbooks | Wrong tenure band read; a date in a footnote breaks the chronology |
Artificial Analysis, which runs its own version of the benchmark, reports that the 100 tasks cover 4,592 source pages and 1,275 atomic grading criteria.[28][29] The Hugging Face dataset card listed 100 examples with 100 unique PDFs and "up to 30" rubric criteria per example, about 468 MB in total. Each criterion carried metadata fields for type, severity, implicitness, subjectiveness and failure mode.[8]
Capability taxonomy
Every item is tagged against eleven capability axes grouped into three tiers.[2]
| Tier | Axes |
|---|---|
| Tier 1: Extraction and grounding | Correctness and completeness; grounding (no facts supplied from prior knowledge); spatial awareness |
| Tier 2: Structural and multimodal comprehension | Semantic reading flow; typographic hierarchy; standard table parsing; chart and multimodal interpretation |
| Tier 3: Advanced reasoning | Complex and multi-page tables; cross-referencing; artifact and noise (scan noise, headers, watermarks, superseded content); unsupported queries |
The "unsupported queries" axis covers items built so that the document lacks the requested information because a figure is redacted, a value was never reported, or a conclusion does not follow. A model gets credit only by saying so plainly. The paper reports that most models "instead produce a fluent, helpful-sounding fabrication, which is scored zero."[2]
Collection workflow
The paper describes six curation steps. A domain expert submits a PDF from their own work with a note and a first-pass answer. Screening models attempt the task, and at least two must commit a major failure, defined as a materially wrong final answer, a dropped piece of decisive evidence or a fabricated claim. Candidates answerable from general priors, dependent on context only the contributor had, or ambiguously worded are then filtered out. The expert writes atomic criteria, including claims the answer must not make. Capability tags and a plain-language note on the parsing challenge are added. Finally, the gold answer has to pass its own rubric and a deliberately bad answer has to fail on the criterion the item was built around.[2] Surge's blog says physicians, attorneys, insurance adjusters, bankers and other professionals wrote and graded the tasks.[9]
The paper's sample items give a sense of the difficulty. One HR task asks for the first five companies to offer paid bereavement leave, and a footnote moves Mastercard's date out of the top five. In a manufacturing task on drilling Ti-8Al-1Mo-1V pipe, the document advises against brad-point bits, yet the paper says every model recommended them. An insurance task turns on an exclusion covering a vehicle used as a residence.[2] The examples on Surge's benchmark page include a fryer wiring-diagram task where the model has to pick the right diagram by serial number, a ceiling-light count on house blueprints, and sound-pressure estimates read from polar diagrams in a NASA technical note.[4]
Grading and evaluation protocol
The paper deliberately avoids ANLS-style text-overlap scoring. Its example is a response that names almost the same companies as the reference but still includes the one company a footnote disqualifies, which an overlap metric would reward.[2] Instead, an LLM judge grades each criterion separately. The paper names the judge as Gemini 3.5 Flash, "calibrated to ensure high agreement with expert human raters." The judge sees neither the source PDF nor the gold answer, only the response and one self-contained criterion. The paper notes that the judge's model family also appears among the evaluated models and argues that criterion-level grading "is far more constrained than the generation task."[2]
In Surge's leaderboard setup, each model gets the question and the PDF, supplied natively as a base64-encoded file input, with no tools and no extra context. That makes each provider's own document handling part of what is measured. Each item is run five times, and the reported number is the strict pass rate averaged over the runs (mean pass@1).[2]
The open-source harness at github.com/surge-ai/gdp-pdf is built on the Inspect AI evaluation framework. Its default judge is google/gemini-3.5-flash. It reports two metrics: all_pass, described as "the headline leaderboard number," and mean_criteria, the average fraction of criteria satisfied. The scorer makes one judge call per criterion and asks for a structured "0" or "1" verdict with a rationale. If the judge output fails to parse five times, that criterion is left unscored and the whole response is excluded from all_pass rather than counted as a failure.[7] The README gives the leaderboard configuration as five epochs per task with Gemini 3.5 Flash as judge, and says the harness can also compute pass@5 and pass^5 (all five runs pass).[7] GitHub shows the repository was created on July 1, 2026.[7]
Public release and "held-out" status
The paper says all 100 items were released publicly on Hugging Face (surgeai/GDP.pdf) with source PDFs, prompts, rubrics and domain labels.[2] The dataset card described the set as "published as a held-out benchmark." In that sense it is evaluation-only: the card provided a single test split, "intentionally no train split," asked that the data be kept out of training corpora, and included a canary GUID for contamination detection. The card says reporting numbers on the benchmark "implies the model was not exposed to it during any training stage."[8] It is therefore not a private test set. Artificial Analysis lists GDP.pdf among the index evaluations that are not private.[29]
The license terms differ between sources. The paper's profile table says the examples are released under Apache 2.0 and that third-party PDFs keep their original rights.[2] The Hugging Face card instead released the prompts, rubrics and metadata under CC BY 4.0. It said the PDFs, collected from public sources, "are not ours and are not covered by the cc-by-4.0 license," were included "for the limited purpose of non-commercial research and evaluation," and would be taken down at a rights holder's request.[8] When this article was checked on September 30, 2026, the surgeai/GDP.pdf page on Hugging Face returned a 404 error, and Surge's Hugging Face organization listed two public datasets, both from its DAYJOB benchmark. The card text quoted here comes from an archived copy of the page dated June 11, 2026.[8]
Separately from the benchmark, Surge sells a "companion" set of GDP.pdf-style training tasks. In a September 29, 2026 post it said it had post-trained Kimi K2.7 on 1,465 of these tasks (one PDF at a time, no tools). According to Surge, the model's strict pass rate on "the held-out benchmark set" rose from 11.0% to 24.5%, moving it from 35th to 9th on the public leaderboard. Surge also reported that on 220 GDPval tasks, the share of tasks scoring at least 90% rose from 19.5% to 32.7%. These results are Surge's own and were not independently replicated in any source found for this article.[11]
Scores
GDP.pdf scores are published by at least four kinds of parties: Surge itself, Artificial Analysis, model developers running their own harnesses, and developers quoting other people's numbers. The tables below keep these sources separate. Figures from different tables should not be ranked together.
Surge's pilot snapshot (April 14, 2026)
The first arXiv version reported a frozen snapshot taken at public release. It covered seven models, all given the PDF natively with no tools, and scored by strict pass rate with 95% Wilson confidence intervals over the 100 items.[3]
| Model | Strict pass rate | 95% CI |
|---|---|---|
| Gemini 3.1 Pro | 15% | 9.3-23.3 |
| Claude Opus 4.7 | 14% | 8.5-22.1 |
| GPT-5.4 | 11% | 6.3-18.6 |
| Grok 4.20 | 7% | 3.4-13.7 |
| Kimi K2.5 | 6% | 2.8-12.5 |
| Mistral Large 3 | 3% | 1.0-8.5 |
| Nova 2 Pro | 1% | 0.2-5.4 |
The authors noted that the confidence intervals "are wide and overlap heavily among the top three models, so the ordering at the top of the table should be read as provisional."[3] Surge's launch post on X likewise said "every frontier model scored under 15%."[6]
Surge leaderboard (July 2026 paper snapshot and September 2026 live board)
Version 3 of the paper reports 17 models as of July 2026, averaged over five runs: GPT-5.6 Sol 30.7%, Claude Fable 5 (Adaptive Max) 29.8%, GPT-5.5 (xHigh) 26%, GPT-5.6 Terra 24.7%, Claude Opus 4.8 (Adaptive Max) 24%, GPT-5.6 Luna 22.7%, Claude Opus 4.7 21%, Claude Sonnet 4.6 18%, Gemini 3.1 Pro 17%, Gemini 3.5 Flash and Grok 4.5 14% each, Kimi K2.6 12%, Gemini 3 Flash 10%, Grok 4.3 8%, and Mistral Large 3, Nova 2 Pro and Nemotron 3 Nano Omni 2% each.[2] Surge's July blog post called GPT-5.6 Sol's 30.7% "the first score above 30% since we released the benchmark in April."[9]
As of September 30, 2026, Surge's live leaderboard listed 39 configurations. The top 15:[4]
| Position | Model (setting as labelled by Surge) | Strict pass rate |
|---|---|---|
| 1 | GPT-6 Astra (Max reasoning) | 34.2% |
| 2 | GPT-5.6 Sol (Max reasoning) | 30.7% |
| 3 | Claude Opus 5.5 (Adaptive/Max) | 30.6% |
| 4 | Claude Fable 5 (Adaptive/Max) | 29.8% |
| 5 (tie) | Claude Fable 5.1 (Adaptive/Max) | 27.6% |
| 5 (tie) | Muse Spark 1.3 (xHigh reasoning) | 27.6% |
| 7 | GPT-6 Sol (Max reasoning) | 26.4% |
| 8 | GPT-5.5 (xHigh reasoning) | 26% |
| 9 | GPT-5.6 Terra (Medium reasoning) | 24.7% |
| 10 (tie) | Claude Opus 4.8 (Adaptive/Max) | 24% |
| 10 (tie) | Claude Opus 5 (Adaptive/Max) | 24% |
| 12 | Gemini 3.7 Flash (High reasoning) | 23.8% |
| 13 (tie) | Gemini 3.8 Flash (High reasoning) | 23.2% |
| 13 (tie) | Qwen 3.8 Max (xHigh reasoning) | 23.2% |
| 15 | GPT-6 Luna (Max reasoning) | 23% |
Lower on the same board were Grok 4.7 (xHigh) at 22.8%, DeepSeek V4.1 Flash (Max) at 19.8%, Kimi K3 (Max) at 19%, Muse Glimmer 30B (xHigh) at 11.8%, and Mistral Large 3, Nemotron 3 Nano Omni and Nova 2 Pro at 2% each. GPT-6.1 Sol, released on September 29, did not yet appear.[4] Epoch AI's benchmarking hub takes its GDP.pdf chart directly from this public Surge leaderboard.[35]
Artificial Analysis implementation
Artificial Analysis added GDP.pdf to its Intelligence Index in version 4.2 (announced September 4, 2026) with a 10% weight in the "General" category. It described the benchmark as testing "single-turn professional document reasoning across 100 PDFs and ten domains."[30][29] Its implementation departs from Surge's in two main ways. Every PDF is converted to text with LiteParse 2.5.0 and English OCR, and page images rendered at 150 DPI are added for models that accept images, instead of using providers' native document inputs. Artificial Analysis argues those inputs "are opaque to the user and sit above the model layer." The judge is GPT-5.6 Luna (medium), which sees the task prompt, the answer and one criterion. Each task is attempted five times, for a fixed denominator of 500 attempts, and errors count as zero. Artificial Analysis states that because "the document input and judge differ," its scores and Surge's "are not directly comparable."[29][28] Version 4.3 of the index (September 7, 2026) also changed GDP.pdf image handling, resizing oversized images in more cases so they can be sent to the model.[29][31]
Artificial Analysis's own figures for the same models have also moved over time, though it has not said which change caused the movement. The v4.2 announcement put GPT-6 Astra at 33.2%, GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%.[30] Artificial Analysis's September 9 write-up of Astra gave 31% for Astra and 27% for GPT-5.6 Sol.[32] When fetched on September 30, 2026, its GDP.pdf leaderboard showed these all-pass rates, among others:[28]
| Model (Artificial Analysis label) | All-pass | Mean criterion pass |
|---|---|---|
| GPT-6 Astra (xhigh) | 32.2% | 81.8% |
| GPT-6.1 Sol (high) | 32.0% | 82.9% |
| GPT-6.1 Sol (xhigh) | 31.8% | 81.4% |
| GPT-6 Astra (max) | 31.0% | 81.8% |
| GPT-6.1 Sol (max) | 31.0% | 81.4% |
| Muse Spark 1.3 (max) | 26.6% | 78.1% |
| Claude Opus 5.5 (max with fallback) | 26.2% | 81.4% |
| Claude Fable 5.1 (max with fallback) | 26.2% | 79.7% |
| Claude Sonnet 5.5 (max with fallback) | 25.8% | 79.0% |
| GPT-6 Sol (max) | 24.8% | 77.5% |
| Qwen3.8 Max (0902) | 22.8% | 76.0% |
| Kimi K3 (max) | 22.0% | 75.7% |
| Gemini 3.8 Flash (high) | 21.0% | 74.5% |
| Grok 4.7 (xhigh) | 20.0% | 70.2% |
| Step 5 Preview | 14.8% | 67.2% |
| Mistral Medium 3.5 | 2.8% | 46.5% |
Mean criterion pass rates for the models above Step 5 Preview fall mostly between 70% and 83%, while their all-pass rates range from 20% to 32%. Models satisfy most individual rubric checks, but they rarely satisfy all of them on a single task. Artificial Analysis said on September 29 that GPT-6.1 Sol made "a 6 point jump in GDP.pdf" over GPT-6 Sol, and on September 22 that Claude Opus 5.5 "remains behind" on GDP.pdf.[33][34]
Same model, different sources
The same model can carry quite different GDP.pdf figures depending on the source. Some examples, with each figure's setting:
| Model | Surge leaderboard | Artificial Analysis | Developer-reported |
|---|---|---|---|
| GPT-5.6 Sol | 30.7% (max)[4] | 28.2% (v4.2 announcement); 27% (Sept 9)[30][32] | 30.7% (OpenAI, GPT-5.6 post)[14]; 40.0% (Google's own run, Gemini 3.8 Flash card)[25] |
| GPT-6 Astra | 34.2% (max)[4] | 32.2% (xhigh), 31.0% (max)[28] | 31.0% max, 32.2% xhigh (OpenAI chart, GPT-6.1 Sol post)[15]; 31.0% (StepFun table)[27] |
| Claude Opus 5 | 24% (Adaptive/Max)[4] | not among the 30 models shown by default | 37.0% (Google's own run)[25]; 21.6% (StepFun table)[27]; 83.4% mean criteria, no tools (Anthropic internal harness)[19] |
| Gemini 3.7 Flash | 23.8% (High)[4] | not among the 30 models shown by default | 34.0% (Google's own run, 3.7 and 3.8 Flash cards)[23][25] |
| Claude Sonnet 5 | not listed | not listed | 28.0% (Google's own run)[23]; 67.5% mean criteria, no tools (Anthropic)[18] |
Several patterns show up in the primary sources:
- OpenAI's GPT-6.1 Sol chart matches Artificial Analysis. The GDP.pdf chart in OpenAI's September 29, 2026 launch post plots five reasoning settings per model. Its data gives GPT-6.1 Sol 27.0% (low), 30.0% (medium), 32.0% (high), 31.8% (xhigh) and 31.0% (max), against GPT-6 Astra's best of 32.2% (xhigh) and a best of 28.8% (high) for a configuration labelled "Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback."[15] Every value that also appears on Artificial Analysis's leaderboard is identical there, including Astra at 32.2% (xhigh) and 31.0% (max), GPT-6.1 Sol at 32.0% (high), GPT-6 Sol at 24.8% (max) and the Opus 5.5 configuration at 26.2% (max).[15][28] The post says that "Evaluations of competitor models were taken from publicly available reports."[15] OpenAI's own figure for Astra is therefore about three points below Astra's 34.2% on Surge's leaderboard. The GPT-6 Astra launch post itself did not include GDP.pdf in its tables.[16]
- Google's figures are largely self-computed. Google's methodology for Gemini 3.7 Flash says the GDP.pdf results "for GPT-5.6 Terra and Muse Spark 1.2 are taken from the official public leaderboard. Results for Gemini models and Sonnet 5 are self computed."[24] That table mixed two measurement methods in a single row: 34.0% for Gemini 3.7 Flash (self-computed) beside 24.7% for GPT-5.6 Terra (Surge's leaderboard).[23] For Gemini 3.8 Flash, Google wrote that "GDP.PDF results for all models are self computed" and labelled the metric "All pass rate". It reported 35.0% for 3.8 Flash, 34.0% for 3.7 Flash, 37.0% for Claude Opus 5, 28.0% for Claude Sonnet 5, 40.0% for GPT-5.6 Sol and 29.0% for GPT-5.6 Terra.[25][26] Every one of those models that is also on Surge's board (all except Sonnet 5) scores higher in Google's table than on Surge's leaderboard.[4]
- StepFun's table matches Artificial Analysis. StepFun's September 2026 Step 5 Preview page gave GDP.pdf scores of 14.8% for Step 5 Preview (High), 11.2% for GLM-5.3 (Max), 22.0% for Kimi K3 (Max), 31.0% for GPT-6 Astra (Max), 26.2% for Claude Fable 5.1 (Max) and 21.6% for Claude Opus 5 (Max).[27] The values that also appear on Artificial Analysis's leaderboard (Step 5 Preview, Kimi K3, GPT-6 Astra max, Fable 5.1) match it.[28] The page does not state a source for the GDP.pdf row.[27]
- Anthropic reports a different metric. See below.
Anthropic's reporting
Anthropic first reported GDP.pdf in the Claude Fable 5 and Mythos 5 system card (June 9, 2026). In that card, Surge ran Claude Fable 5 with its standard harness: adaptive thinking, max effort, no tools, full 100 prompts. The card reports a strict pass rate of 29.8% for Fable 5, against 22.5% for Claude Opus 4.8, 24.9% for GPT-5.5 and 16.7% for Gemini 3.1 Pro. It describes Surge's judge as Gemini 3 Flash, whereas the paper and harness name Gemini 3.5 Flash.[17][2][7] Anthropic also ran GDP.pdf on an internal harness with and without tools. That harness truncates, rather than drops, PDFs larger than its 32 MB API request limit, and in the tools setting gives the model a container with Python and an image-cropping tool. It reports the mean criteria pass rate (72.7% without tools and 87.6% with tools for Claude Mythos 5). The card adds: "We note that we were not able to reproduce Surge's reported numbers and that both mean criteria pass rates and strict pass rates trail below those from Surge's runs."[17]
Later cards kept the internal harness and changed its settings. The Claude Sonnet 5 card (June 30) switched the judge to Claude Opus 4.7, raised the output limit to 128k tokens, and reported 67.5% without tools and 81.6% with tools for Sonnet 5. It repeated that Surge's numbers could not be reproduced.[18] The Claude Opus 5 card (July 24) disclosed "a bug in our harness which was unnecessarily truncating PDFs longer than 100 pages but smaller than our 32MB request size limit." After fixing it and re-running earlier models, Anthropic reported Opus 5 at 83.4% and 85.5%, Opus 4.8 at 77.5% and 84.8%, and Mythos 5 at 81.8% and 87.3% (no tools and with tools respectively). The revised Mythos 5 no-tools figure is 9.1 points above the 72.7% in the June card.[19] The Claude Fable 5.1 card (September 1) reported 85.4% without tools and 85.1% with tools.[20] The Claude Opus 5.5 (September 22) and Claude Sonnet 5.5 (September 28) system cards cover multimodal evaluation with Chartography, another Surge benchmark, and do not report GDP.pdf.[21][22]
Role in lab marketing
Surge has promoted each adoption of the benchmark. Its blog posts are titled "Anthropic cited GDP.pdf and Riemann-bench in their Fable 5 and Mythos 5 system card" and "OpenAI cites GDP.pdf in its GPT-5.6 release."[10][9] The first argues that "the easy benchmarks are saturated" and that the evaluations still separating frontier models "are the ones built by people with deep domain expertise."[10] The second notes that GPT-5.6 Sol scored 94.6% on GPQA Diamond and 90.4% on BrowseComp, but 30.7% on GDP.pdf, "one of the lowest capability benchmark scores OpenAI reported for its flagship model."[9] In OpenAI's GPT-5.6 post, GDP.pdf ("gdp.pdf") appears in the multimodal table beside MMMU Pro, with 30.7% for GPT-5.6 Sol, 24.7% for Terra, 22.7% for Luna, 26% for GPT-5.5, 29.8% for Claude Fable 5, 22.5% for Claude Opus 4.8 and 16.7% for Gemini 3.1 Pro Preview.[14]
In the GPT-6.1 Sol post, GDP.pdf is one of six charts that plot score against cost per task. OpenAI's claims are that GPT-6.1 Sol "scores higher than Opus 5.5 with fallbacks at less than half the cost per task across the tested reasoning settings" and "approaches GPT-6 Astra's state-of-the-art performance at roughly one-fifth the cost per task."[15] In the chart data, GPT-6.1 Sol's cost per task runs from about $0.33 (low) to $0.42 (max), GPT-6 Astra's from about $1.70 to $2.08, and the Opus 5.5 configuration's from about $0.76 to $1.55.[15]
Surge also uses GDP.pdf in its own composite. The Tuesday Work Index, introduced on August 18, 2026, combines eight Surge benchmarks and lists GDP.pdf as its "professional multimodal reasoning" component. The launch post opens by noting that "the best model in the world scores 30.7% on GDP.pdf."[12] Surge's September 10 write-up on that index reported that Muse Spark 1.3 rose from 16.0% to 27.6% on GDP.pdf at xHigh, and that Gemini 3.8 Flash scored 23.4% at Medium against 23.2% at High, one of several cases where more reasoning did not help.[13]
Limitations and criticism
- Adversarial selection. An item entered the benchmark only if at least two frontier models failed it at collection time. The paper itself cautions that absolute pass rates therefore "measure performance on adversarially selected professional tasks, not on professional document work at large."[2] Selection also depended on which screening models were available at the time. The paper says the screening pool was varied so that selection "is not tied to the failure profile of any single model or family."[2]
- Small sample. With 100 items, a single task is worth one percentage point of strict pass rate. The April pilot's 95% intervals spanned about 14 points for the top models.[3] Later leaderboard versions average five runs but do not publish intervals.[2][4]
- Dependence on the harness. Surge's setup measures each provider's native PDF handling. The paper says this is intentional ("each provider's own document handling is therefore part of what is measured"), but it means scores mix model ability with API document pipelines.[2] Artificial Analysis rejected native document inputs for that reason and uses its own text-plus-image pipeline.[29] Anthropic could not reproduce Surge's figures, and its own harness had a PDF-truncation bug that depressed scores until July 2026.[17][19] In the Gemini 3.8 Flash card, Google's self-computed scores are higher than Surge's for every model that also appears on Surge's board.[25][4]
- LLM judging. Criteria are graded by an LLM, and the paper acknowledges that the judge's model family (Gemini) is among those evaluated.[2] The judge also varies: Gemini 3.5 Flash for Surge, GPT-5.6 Luna for Artificial Analysis, Claude Opus 4.7 in Anthropic's internal runs from June 30, 2026 onward, and DeepSeek V4 Pro in one vendor's study.[2][29][18][36]
- Moving headline numbers. Surge's own messaging has changed as the leaderboard moved. The April 14 announcement said every frontier model scored under 15%. The current launch blog says no model "cleared 25%" and also that "Every frontier model scored under 30%." The live leaderboard's top score is 34.2%.[6][5][4] Artificial Analysis scores changed between index versions 4.2 and 4.3 after the image-handling change.[30][28][31]
- Public data and contamination. The full task set, including rubrics, was released publicly. The dataset card relies on a canary string and a request not to train on the data.[8] Surge also sells GDP.pdf-style training tasks and reports large gains from training on them, while describing the benchmark set as held out.[11]
- Evidence-layer effects. Pulse, a document-extraction vendor, reported in April 2026 that feeding models its structured extraction alongside the PDF raised the average whole-task success of three models (GPT-5.5, Claude Opus 4.8 and Gemini 3.1) from 21.67% with the PDF alone to 31.67%. Pulse graded its own runs with DeepSeek V4 Pro. Its 21.67% baseline equals the average of Surge's published scores for those three models (25%, 23% and 17%), so the comparison sets Pulse-judged runs against Surge-judged ones.[36][5] This is a vendor's own result, but it is consistent with the paper's view that many failures come from structure lost before reasoning begins.[2]
Findings on model failures
The paper's error analysis groups failures into recurring patterns:[2]
| Failure pattern | Example from the paper |
|---|---|
| Table misalignment | On an HR benefits item, one model read the wrong tenure band; another read the right band and reported numbers not in it |
| Chart and figure misreads | Models found the right threshold-irradiance plot on a datasheet but returned implausible values |
| Dropped footnotes, cross-references and exclusions | Most models ignored the footnote that reorders the bereavement-leave ranking; insurance answers stopped before the exclusions |
| Priors overriding the document | Every model recommended brad-point bits for Ti-8Al-1Mo-1V pipe although the document advises against them |
| Spatial reasoning | Fixture and window counts on floor plans were wrong because symbols were not matched to legends or schedules |
| Noise and supersession | Models quoted deed sections already replaced by amendments; a "PENDING" registration label was treated as undermining sale prices |
The authors conclude that longer context windows alone may not close the gap. They argue that page representations which keep structure intact, better chart and table parsing, retrieval that preserves visual evidence, and better uncertainty calibration for abstention items would help more.[2] They call GDP.pdf "a benchmark for the floor, not just the ceiling," and say the 2 to 30.7% range in their July table "is not a number that supports unsupervised use."[2]
See also
References
- ^1 ^2 ^3 ^4 ^5arXiv. "GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents" (abstract page, arXiv:2607.11192; v1 July 13, 2026, v3 July 15, 2026). Suhaas Garre, Emily Ritchie, Sushant Mehta, Edwin Chen. arxiv.org/...2607.11192
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30Garre, S., Ritchie, E., Mehta, S., Chen, E. "GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents." arXiv:2607.11192v3, July 15, 2026. arxiv.org/...2607.11192v3
- ^1 ^2 ^3 ^4 ^5 ^6Garre, S., Ritchie, E., Mehta, S., Chen, E. "GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents." arXiv:2607.11192v1 (pilot evaluation), July 13, 2026. arxiv.org/...2607.11192v1
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13Surge AI. "GDP.pdf Benchmark" (leaderboard and example tasks). Accessed September 30, 2026. surgehq.ai/...gdp-pdf
- ^1 ^2 ^3 ^4Surge AI. "GDP.pdf Benchmark: Can Frontier Models Master the Documents that Run the World?" Surge AI blog, April 14, 2026 (page also shows September 16, 2026). surgehq.ai/...ter-the-documents-that-run-the-world
- ^1 ^2 ^3Surge AI (@HelloSurgeAI). "Introducing GDP.pdf: an expert multimodal reasoning benchmark for the documents that run the world." X, April 14, 2026. x.com/...2044115915718177208
- ^1 ^2 ^3 ^4Surge AI. "surge-ai/gdp-pdf" (GDP.pdf Harness; README and src/gdp_pdf/scorer.py). GitHub, created July 1, 2026. github.com/...gdp-pdf
- ^1 ^2 ^3 ^4 ^5Surge AI. "surgeai/GDP.pdf" dataset card. Hugging Face, archived June 11, 2026. web.archive.org/...GDP.pdf
- ^1 ^2 ^3 ^4Surge AI. "OpenAI cites GDP.pdf in its GPT-5.6 release." Surge AI blog, July 9, 2026. surgehq.ai/...openai-gpt-5-6-gdp-pdf-benchmark
- ^1 ^2Surge AI. "Anthropic cited GDP.pdf and Riemann-bench in their Fable 5 and Mythos 5 system card." Surge AI blog, June 2026. surgehq.ai/...le-5-mythos-5-cites-surge-benchmarks
- ^1 ^2Surge AI. "What Does GDP.pdf Teach, Beyond PDFs?" Surge AI blog, September 29, 2026. surgehq.ai/...gdp-pdf-post-training
- ^Surge AI. "Introducing the Tuesday Work Index: Can AI Get Through an Ordinary Day?" Surge AI blog, August 18, 2026. surgehq.ai/...tuesday-frontier-work-index
- ^Surge AI. "Fable 5.1, Muse Spark 1.3, and Gemini 3.8 Flash on the Tuesday Work Index." Surge AI blog, September 10, 2026. surgehq.ai/...fable-5-1-tuesday-work-index
- ^1 ^2 ^3OpenAI. "GPT-5.6: Frontier intelligence that scales with your ambition." July 2026. openai.com/...gpt-5-6
- ^1 ^2 ^3 ^4 ^5 ^6 ^7OpenAI. "Introducing GPT-6.1 Sol." September 29, 2026 (chart data read from the rendered page). openai.com/...introducing-gpt-6-1-sol
- ^OpenAI. "GPT-6 Astra: A new generation of intelligence." September 2026. openai.com/...gpt-6-astra
- ^1 ^2 ^3 ^4Anthropic. "System Card: Claude Fable 5 & Claude Mythos 5." June 9, 2026, section 8.16.1. www-cdn.anthropic.com/...6dd7cb2e3c342ee809620.pdf
- ^1 ^2 ^3 ^4Anthropic. "System Card: Claude Sonnet 5." June 30, 2026, section 8.10.1. anthropic.com/claude-sonnet-5-system-card
- ^1 ^2 ^3 ^4Anthropic. "System Card: Claude Opus 5." July 24, 2026, section 8.12.4. www-cdn.anthropic.com/...s%205%20System%20Card.pdf
- ^1 ^2Anthropic. "System Card: Claude Fable 5.1 & Claude Mythos 5.1." September 1, 2026, section 8.14.4. www-cdn.anthropic.com/...205.1%20System%20Card.pdf
- ^Anthropic. "System Card: Claude Opus 5.5." September 22, 2026. www-cdn.anthropic.com/...205.5%20System%20Card.pdf
- ^Anthropic. "System Card: Claude Sonnet 5.5." September 28, 2026. www-cdn.anthropic.com/...205.5%20System%20Card.pdf
- ^1 ^2 ^3 ^4Google DeepMind. "Gemini 3.7 Flash Model Card." August 13, 2026. deepmind.google/...gemini-3-7-flash
- ^Google DeepMind. "Gemini 3.7 Flash: Model evaluation, approach, methodology & results." August 2026. deepmind.google/...gemini-3-7-flash
- ^1 ^2 ^3 ^4 ^5 ^6Google DeepMind. "Gemini 3.8 Flash Model Card." September 2, 2026. deepmind.google/...gemini-3-8-flash
- ^Google DeepMind. "Gemini 3.8 Flash: Model evaluation, approach, methodology & results." September 2026. deepmind.google/...gemini-3-8-flash
- ^1 ^2 ^3 ^4StepFun. "Step 5 Preview." September 2026. stepfun.com/step-5-preview
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Artificial Analysis. "GDP.pdf Benchmark Leaderboard." Accessed September 30, 2026. artificialanalysis.ai/...gdp-pdf
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Artificial Analysis. "Intelligence Benchmarking Methodology" (GDP.pdf section and changelog). Accessed September 30, 2026. artificialanalysis.ai/...intelligence-benchmarking
- ^1 ^2 ^3 ^4 ^5Artificial Analysis. "Announcing Artificial Analysis Intelligence Index v4.2." September 4, 2026. artificialanalysis.ai/...s-intelligence-index-v4-2
- ^1 ^2Artificial Analysis. "Announcing the Artificial Analysis Intelligence Index v4.3." September 7, 2026. artificialanalysis.ai/...s-intelligence-index-v4-3
- ^1 ^2Artificial Analysis. "Benchmarking GPT-6 Astra." September 9, 2026. artificialanalysis.ai/...benchmarking-gpt-6-astra
- ^Artificial Analysis. "GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence." September 29, 2026. artificialanalysis.ai/...h-near-astra-intelligence
- ^Artificial Analysis. "Claude Opus 5.5." September 22, 2026. artificialanalysis.ai/...claude-opus-5-5
- ^Epoch AI. "GDP.pdf" (benchmarking hub, methodology). Accessed September 30, 2026. epoch.ai/...gdp-pdf
- ^1 ^2Pulse AI (Sid and Ritvik). "GDP.pdf and the Retrieval Layer: How Pulse Improves Enterprise Document AI." April 29, 2026. runpulse.com/...gdp-pdf-and-the-retrieval-layer
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 5,831 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification V5 (xg14, 30 Sep 2026): ~120 claims vs arXiv paper, Surge board, AA, OpenAI chart data, Google/StepFun tables, Anthropic cards, HF archive; 0 material, 2 minor fixed
Cite this page: AI Wiki. "GDP.pdf." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/gdp_pdf