Chartography
Chartography is a benchmark for professional chart understanding created by Surge AI. It contains 100 tasks, each pairing a chart drawn from professional practice (Kaplan-Meier survival curves, candlestick charts, contour maps, wind roses, Sankey diagrams, Bode plots, three-dimensional surface plots and similar domain-specific formats) with a question written by a practitioner who reads such charts for a living. Visually estimated answers are graded against acceptable ranges that the expert author set for each chart, rather than against a single numeric tolerance.[1][2] Surge released the benchmark on 16 July 2026 with a public leaderboard, an evaluation harness on GitHub and an accompanying paper, which was posted to arXiv on 11 August 2026 and accepted at an ECCV 2026 workshop.[1][2][4]
At launch the best of 30 frontier-model configurations, GPT-5.6 Sol at maximum reasoning effort, passed 45.0% of graded trials, while the same model families score around 80-90% on older chart benchmarks such as ChartQA, CharXiv and ChartMuseum.[1] Chartography has since been adopted by model developers as a multimodal evaluation. Anthropic has reported it in the system cards for Claude Opus 5 (July 2026), Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5, under both "no tools" and "with tools" settings, and DeepSeek has reported it for two of its V4-series models.[5][6][8][10][11][12] In the Claude Sonnet 5.5 launch post (28 September 2026), Anthropic listed it as "Visual chart recognition" and credited Surge AI as the source of its GPT-6 Sol figure.[9]
| Field | Detail |
|---|---|
| Type | Multimodal benchmark: professional chart reading and reasoning |
| Creator | Surge AI evaluations team[4] |
| Paper authors | Suhaas Garre, Chris Mutty, Sushant Mehta, Edwin Chen (all Surge AI)[1] |
| Released | 16 July 2026 (blog post, leaderboard and GitHub repository)[1][2][4] |
| Paper | arXiv:2608.10677, submitted 11 August 2026; accepted at the 2nd Workshop on Benchmarking Evidence-Aligned Multimodal Reasoning (BEAM 2), ECCV 2026[1] |
| Size | 100 tasks (100 task-image records, 98 byte-distinct images) across 12 domain labels[1] |
| Answer types | 51 range answers, 25 exact numeric, 24 categorical[1] |
| Metric | Mean pass@1 over repeated trials; multi-part answers all-or-nothing[1] |
| Grader | LLM judge (Gemini 3.5 Flash by default) that sees the question, golden answer and response but never the chart[1][4] |
| Official protocol | Single user message, chart inline, no system prompt, no tools[1] |
| Harness | Built on the Inspect AI framework[1][4] |
| Leaderboard | surgehq.ai/benchmarks/chartography[3] |
Name and motivation
Surge says the name combines "chart" and "cartography": "A chart maps a body of data, and reading that map accurately is a professional skill."[2] The benchmark's premise is that in medicine, engineering, finance, manufacturing and the sciences the working source of quantitative information is often a chart rather than a table. The paper's examples are an oncologist weighing treatment options from Kaplan-Meier curves, a structural engineer taking a design wind speed from a contour map, and a trader reading price history from candlesticks. In each case the needed values are not printed on the figure and must be interpolated, traced or projected.[1]
The authors argue that earlier chart benchmarks measure little of this. In their account ChartQA was "largely solved by the Claude 3.5 Sonnet generation of models," frontier models score around 90% on CharXiv and above 80% on ChartMuseum, and many remaining errors trace to defects in ground truth. They describe those datasets as dominated by bar, line, pie and scatter charts, one- or two-step questions, constrained answer formats such as multiple choice, and grading anchored to printed labels or uniform percentage tolerances.[1][2] The paper positions Chartography as a single-chart complement to Surge's GDP.pdf benchmark, which covers reasoning over whole professional PDF documents.[1]
Construction
Surge describes four design principles: charts from real professional domains, questions written by real professionals, visual reasoning beyond label lookup, and grading calibrated to each chart.[1][2]
- Expert authoring. Each task was written by a working professional in the relevant field (the paper lists clinicians, engineers, scientists and financial professionals). Every submission had to include the intended answer, a complete walkthrough of the visual and computational steps, and, for visually estimated quantities, an acceptable range. Open-ended questions were not allowed, questions could assume specialist knowledge but not data beyond the chart, and a question answerable from the prompt alone was not allowed.[1]
- Pilot rounds. Two pilot rounds of 25 submissions each led to revised authoring instructions and a required "answer logic" field showing the author's calculations.[1]
- Independent review. Each task was verified by three additional domain experts, who checked the intended reading, revised unclear wording and checked the acceptable range. Tasks were revised or rejected when the prompt, figure or target did not support a single defensible grading decision.[1]
- Adversarial difficulty screening. Candidate tasks were run against frontier models during curation, and a task was kept only if it produced a meaningful failure in at least one of them.[1] The GitHub README states that "all 100 tasks defeated at least one of two frontier models."[4]
The author walkthroughs were used for construction and verification but are not released as public fields.[1]
Dataset composition
Of the 100 charts, 81 were sourced online and 19 were created by task authors. The most common online sources are Wikimedia projects (15 charts), NIST publications (9) and the CDC (6), alongside NASA, the Bureau of Transportation Statistics and other public repositories.[1]
| Domain label | Tasks | Online-sourced | Expert-created | With explicit range |
|---|---|---|---|---|
| STEM - General | 19 | 15 | 4 | 13 |
| Manufacturing / Supply Chain | 14 | 12 | 2 | 2 |
| STEM - General Engineering | 11 | 10 | 1 | 6 |
| STEM - Mechanical Engineering | 10 | 8 | 2 | 7 |
| Finance / Investing | 10 | 10 | 0 | 4 |
| Healthcare | 10 | 8 | 2 | 4 |
| STEM - Electrical Engineering | 9 | 6 | 3 | 5 |
| STEM - Chemistry | 6 | 2 | 4 | 3 |
| STEM - Civil/Environmental Engineering | 4 | 3 | 1 | 4 |
| STEM - Geosciences | 4 | 4 | 0 | 2 |
| STEM - Biology | 2 | 2 | 0 | 0 |
| STEM - Physics | 1 | 1 | 0 | 1 |
| Total | 100 | 81 | 19 | 51 |
Source: Chartography paper, Table 2.[1]
The release profile in the paper lists a median prompt length of 33.5 words (range 5-163), a median answer length of 7 words (range 1-78), 31 prompts that specify rounding, and images in PNG (80), JPG (18) and WebP (2) format, 80 landscape and 20 portrait, with a median width of 990 pixels (range 310-4,721).[1] Each released row has the fields task_id, prompt, golden_answer, chart_path, domain_combined, source_type and source_url.[1][4]
Chart formats named by Surge include Kaplan-Meier curves, pediatric growth charts, candlestick and financial market charts, contour maps, wind roses, Sankey and material-flow diagrams, Bode plots and engineering response curves, atmospheric back-trajectories, failure-criterion and design charts, control charts, phase diagrams and three-dimensional surface plots. Chart type is not a task-level annotation in the release, so the paper calls this list illustrative rather than a set of strata.[1][2] The operations the tasks demand include estimating values along sparsely labeled axes, interpolating between contours and curves, following thin or overlapping traces, mapping legend entries to the correct marks, and interpreting projected geometry, often several within one task.[1]
Example tasks
Surge published three example tasks with their golden answers:[1][2][3]
| Chart | Question | Golden answer |
|---|---|---|
| Engineering design chart (USDA-SCS) | "For an outlet pipe with a diameter of 48 inches and a flow of 73 cubic feet per second, what is the minimum downstream width of the riprap apron?" | 30 feet (accepted 29-31) |
| Pediatric growth chart (CDC) | "What height corresponds to the 50th percentile for a four-year-old girl? Round to the nearest 0.5 centimeters." | Accepted range 100-101 cm (the paper gives 101.0 cm; the blog post gives 100.5 cm) |
| Wind rose (USDA-ARS) | "In which three directions did winds above 8.5 m/s occur for the longest total duration?" | S, SSW, W |
In the riprap task the flow lies below the plotted range of the 48-inch curve, so professional practice is to read the curve's minimum (26 feet) and add the pipe diameter (4 feet). A representative failure shown on the leaderboard page extrapolated to about 32 feet, invented a formula that tripled the diameter, and answered 44 feet. In the growth-chart task, a model identified as Opus on the leaderboard page followed the correct 50th-percentile curve but answered 101.5 cm. In the wind-rose task, a model identified as Gemini answered S, SSW, SW, omitting W.[1][3]
Grading and metric
Of the 100 golden answers, 51 carry at least one expert-set acceptable range (21 carry more than one, for multi-quantity questions), 25 are exact numeric values, and 24 are categorical (an item, a set or a ranking).[1] The paper gives the rationale with two cases: a wind-speed contour map may support an answer of 150 mph only to within 140-150 mph, while a pediatric growth chart with fine gridlines supports a reading to within 0.5 cm, so no single percentage tolerance fits both.[1]
Grading uses an LLM judge, by default Gemini 3.5 Flash, in one model-graded call per trial. The judge receives the question, the golden answer and the model's response but never the chart, so it rules on answer equivalence rather than re-solving the task. Its rules are:[1][4]
- grade only the model's last committed answer;
- a numeric answer is correct if it falls inside the expert range (inclusive), or matches the golden answer exactly (allowing trivial rounding) where no range exists;
- a categorical answer must match the golden set exactly, with order mattering only when the prompt asks for a ranking;
- multi-part answers are all-or-nothing;
- declaring a question unanswerable is wrong unless the golden answer is itself "cannot be determined."
The judge returns a binary score and rationale per criterion, and a trial passes only if every criterion scores 1. Because Gemini 3.5 Flash is also an evaluated model, the paper notes that its own leaderboard row is a self-judging condition.[1]
The reported score is mean pass@1: the fraction of all graded trials passed. The harness can also compute pass@k and pass^k.[1][4]
Evaluation protocol
Under the official protocol each model receives a single user message with the task prompt followed by the chart image, with no system prompt, no few-shot examples, no image normalization, and no tools (no code execution, cropping or zooming). The model answers free-form. Temperature, maximum output length, provider-side image processing and reasoning controls are treated as properties of the run configuration rather than of the task.[1]
The paper's launch evaluation ran 20 trials per task, or 2,000 graded trials per configuration, and reports that no trial exhausted the judge's four retries.[1] The GitHub README, by contrast, says that matching "the configuration used in the leaderboard" means running the full task set "with ten runs per task and Gemini 3.5 Flash as the judge."[4] The harness is built on Inspect AI, the evaluation framework published by the UK AI Security Institute.[1][4]
In its discussion section, the paper encourages developers of tool-augmented and agentic systems to report results under the same no-tools protocol alongside their augmented configurations.[1]
Launch results
The launch leaderboard, a frozen snapshot dated 16 July 2026, covers 30 configurations of 18 models from eight developers, including a default and a maximum-reasoning configuration where available.[1]
| Rank | Model (configuration) | Developer | Pass@1 (%) ± 95% CI |
|---|---|---|---|
| 1 | GPT-5.6 Sol (max reasoning) | OpenAI | 45.0 ± 1.1 |
| 2 | GPT-5.6 Sol (default) | OpenAI | 39.5 ± 1.2 |
| 3 | Gemini 3.5 Flash | 35.9 ± 1.3 | |
| 4 | Claude Fable 5 (adaptive/max) | Anthropic | 34.8 ± 1.0 |
| 5 | GPT-5.6 Terra (max reasoning) | OpenAI | 34.0 ± 1.2 |
| 6 | GPT-5.6 Luna (max reasoning) | OpenAI | 31.4 ± 1.2 |
| 7 | GPT-5.5 (xHigh reasoning) | OpenAI | 31.0 ± 1.2 |
| 8 | GPT-5.4 (xHigh reasoning) | OpenAI | 29.6 ± 1.1 |
| 9 | Claude Fable 5 (default) | Anthropic | 29.5 ± 1.1 |
| 10 | GPT-5.5 (default) | OpenAI | 28.5 ± 1.3 |
| 11 | GPT-5.6 Terra (default) | OpenAI | 26.6 ± 1.3 |
| 12 | Gemini 3.1 Pro | 26.1 ± 1.3 | |
| 13 | Muse Spark 1.1 (xHigh reasoning) | Meta | 24.4 ± 1.4 |
| 14 | Muse Spark 1.1 (default) | Meta | 23.6 ± 1.4 |
| 15 | GPT-5.6 Luna (default) | OpenAI | 21.4 ± 1.2 |
| 16 | Grok 4.5 (default) | xAI | 17.3 ± 1.1 |
| 17 | Grok 4.5 (high reasoning) | xAI | 16.7 ± 1.1 |
| 18 | Claude Sonnet 5 (adaptive/max) | Anthropic | 16.6 ± 1.1 |
| 19 | Claude Opus 4.7 (adaptive/max) | Anthropic | 16.5 ± 1.0 |
| 20 | Claude Opus 4.8 (adaptive/max) | Anthropic | 15.9 ± 1.0 |
| 21 | Qwen 3.5 Plus | Alibaba | 15.9 ± 1.2 |
| 22 | GPT-5.4 (default) | OpenAI | 14.1 ± 0.9 |
| 23 | Claude Opus 4.7 (default) | Anthropic | 13.6 ± 0.9 |
| 24 | Grok 4.3 (high reasoning) | xAI | 12.9 ± 1.1 |
| 25 | Kimi K2.5 | Moonshot AI | 12.6 ± 1.1 |
| 26 | Claude Sonnet 5 (default) | Anthropic | 12.2 ± 1.1 |
| 27 | Kimi K2.6 | Moonshot AI | 12.2 ± 1.1 |
| 28 | Grok 4.3 (default) | xAI | 11.7 ± 1.0 |
| 29 | Claude Opus 4.8 (default) | Anthropic | 11.3 ± 1.0 |
| 30 | Mistral Large 3 | Mistral | 9.0 ± 1.0 |
Source: Chartography paper, Table 3 (no tools, 100 tasks x 20 trials).[1] The paper cautions that entries with overlapping 95% confidence intervals should not be read as holding distinct ranks.
Reasoning effort. Of 12 model families tested at both a default and an elevated setting, 11 scored higher at the elevated setting. The median gain was 4.5 percentage points, ranging from +15.5 points (GPT-5.4) and +10.0 (GPT-5.6 Luna) down to +0.8 (Muse Spark 1.1) and -0.6 (Grok 4.5, where high reasoning scored below default). The authors read the spread as a sign that when the binding constraint is visual perception, extra deliberation over a wrong reading cannot repair it.[1]
Test-time compute. Plotting score against cost and generated tokens per trial, the paper finds the cost frontier steep at the low end (Gemini 3.5 Flash reaches 35.9% at a fraction of most competitors' cost) and flat at the top, where going from roughly 40% to 45% multiplies cost per trial several-fold. Some maximum-reasoning configurations emit 20,000 to 30,000 or more tokens per trial yet trail cheaper configurations. In trace review, longer reasoning "often preserves an incorrect visual observation rather than correcting it."[1] A later Surge post found the same pattern within one model: Gemini 3.8 Flash scored 42.5% at Medium reasoning with about 7,700 tokens per response and 40.9% at High with about 26,800.[14]
Failure modes
The paper groups failed trajectories into five recurring categories:[1][2]
- Missed visual elements. Thin, faint or overlapping features drop out of the model's reading. In one Sankey task every evaluated model found the largest flow into a sector and missed the three thinner flows beside it, and every model failed the task.
- Right feature, incorrect value. Models pick the correct curve or contour and then misjudge its position: reading a peak two gridlines too low, assigning a location to the neighboring contour band, or misreading an intersection. Suggested mechanisms include choosing the wrong tick interval, treating a nonlinear scale as linear, snapping to a nearby printed label, and reporting unjustified precision.
- Complex geometry. Three-dimensional surface plots produced some of the lowest scores. According to the paper, no model used the projected gridlines that make such plots readable to trained humans.
- Domain conventions. Charts encode rules never stated in the prompt, such as not extrapolating beyond a plotted range. In the riprap task, the paper says, every model missed the rule.
- Compounding errors. A model chooses the correct curve, formula and procedure but misreads one input, and the final answer lands outside the range. Premature rounding of intermediate reads makes this worse.
From these the authors conclude that models succeed at semantic recognition (knowing what a chart shows and which curve matters) but fail at "metric grounding," anchoring an answer to the specific mark or axis position that determines it, and that perception rather than reasoning is often the bottleneck.[1]
Live leaderboard
Surge said it would maintain the public leaderboard as new models launch, and it has added entries since the July snapshot.[1] The leaderboard page, as fetched on 29 September 2026, listed 42 entries. Selected rows:[3]
| Model (configuration shown) | Score |
|---|---|
| GPT-6 Astra (Max reasoning) | 71% |
| Claude Opus 5.5 (Adaptive/Max) | 66.3% |
| GPT-6 Sol (Max reasoning) | 53.6% |
| Claude Fable 5.1 (Adaptive/Max) | 46.2% |
| GPT-5.6 Sol (Max reasoning) | 45% |
| Gemini 3.8 Flash (High reasoning) | 40.9% |
| Gemini 3.7 Flash (High reasoning) | 40.4% |
| Gemini 3.5 Flash (Medium reasoning) | 35.9% |
| Claude Fable 5 (Adaptive/Max) | 34.8% |
| Qwen 3.8 Max (xHigh reasoning) | 29.1% |
| Claude Opus 5 (Adaptive/Max) | 27.3% |
| Kimi K3 (Max reasoning) | 26.6% |
| Grok 4.7 (xHigh reasoning) | 14.7% |
| DeepSeek V4.1-Flash (Max reasoning) | 11.6% |
| DeepSeek V4-Flash Vision (experimental) (Max reasoning) | 11.2% |
| Mistral Large 3 | 9% |
Source: Surge AI leaderboard page, fetched 29 September 2026.[3] Anthropic's system cards describe Surge's published figures as "evaluated without tools."[6][8][10]
Use in model releases
No tools versus with tools
Surge's own protocol is tool-free. Developers who report Chartography in release materials usually add a tool-augmented setting, and the two settings produce very different numbers. In Anthropic's evaluations the "with tools" model gets a container holding the image file and standard libraries, plus an image-cropping tool.[5][6][8][10] DeepSeek ran its "w/ tools" evaluation of V4.1-Flash inside a Claude Code harness with a 512K-token context window.[12] Anthropic's with-tools figures run from 75.0% (Claude Opus 4.8) to 90.2% (Claude Sonnet 5.5), while its no-tools figures run from 15.6% (Claude Sonnet 5) to 64.4% (Claude Opus 5.5).[5][9][10] The same split explains why the Claude Opus 5.5 launch post (which shows with-tools scores) lists Opus 5.5 at 89.0%, while the Sonnet 5.5 launch post (which shows no-tools scores) lists it at 64.4%.[7][9]
Anthropic has reported that tool use is a more cost-effective way to raise scores than extra thinking alone. The Opus 5 system card says that "leveraging the models' agentic coding capabilities to manipulate, analyze, and crop images can be significantly more cost-effective than simply enabling adaptive thinking," and the Fable 5.1 card repeats the finding.[5][6]
Reported results
| Source (date) | Model | No tools | With tools | Notes |
|---|---|---|---|---|
| Claude Opus 5 system card (24 Jul 2026)[5] | Claude Opus 5 | 29.6% | 83.0% | Adaptive thinking, max effort, five runs |
| Claude Mythos 5 | 36.0% | 85.2% | ||
| Claude Opus 4.8 | 17.0% | 75.0% | ||
| DeepSeek V4-Flash-Vision-Exp release (21 Aug 2026)[11] | DeepSeek V4-Flash-Vision-Exp | 64.3; setting not labeled, listed under "Multimodal Agent Evaluation" | ||
| Opus-4.8 (as reported by DeepSeek) | 65.0 | |||
| Claude Fable 5.1 system card (1 Sep 2026)[6] | Claude Fable 5.1 | 42.6% | 86.2% | Section 8.14.1 |
| Claude Fable 5 | 36.6% | 84.2% | ||
| Claude Opus 5 | 29.6% | 83.0% | ||
| DeepSeek V4.1-Flash model card (Sep 2026)[12] | DeepSeek V4.1-Flash | 78.9 | Pass@1, max effort, Claude Code harness | |
| Opus-5.0 (as reported by DeepSeek) | 84.0 | |||
| GPT-5.6 Sol (as reported by DeepSeek) | 79.9 | |||
| Kimi K3 (as reported by DeepSeek) | 68.1 | |||
| Claude Opus 5.5 system card and launch post (22 Sep 2026)[7][8] | Claude Opus 5.5 | 64.4% | 89.0% | Gemini 3.5 Flash judge for all models |
| Claude Fable 5.1 | 44.8% | 88.4% | ||
| Claude Opus 5 | 29.8% | 83.4% | ||
| Claude Sonnet 5.5 system card and launch post (28 Sep 2026)[9][10] | Claude Sonnet 5.5 | 61.6% | 90.2% | Launch post shows no-tools scores only |
| Claude Sonnet 5 | 15.6% | From the launch post | ||
| Claude Opus 5.5 | 64.4% | 89.0% | ||
| Claude Fable 5.1 | 44.8% | 88.4% | ||
| GPT-6 Sol | 53.6% | Surge AI figure (see note below) |
In the Sonnet 5.5 system card, Anthropic places Sonnet 5.5's no-tools score below Opus 5.5 and between GPT-6 Sol and GPT-6 Astra, with the GPT-6 figures "as publicly reported by Surge AI," and says that with tools Sonnet 5.5 "is competitive with Claude Opus 5.5."[10] Footnote 4 of the Sonnet 5.5 launch post notes that OpenAI had recently fixed a bug that degraded image understanding in GPT-6 Sol and that "Chartography scores from Surge AI" might not yet reflect the latest version of the model; it adds that Anthropic's "internal testing of Chartography suggests its score was not impacted."[9]
Comparability
Scores for the same model differ across sources, and the documents explain some of the reasons:
- Judge model. The Opus 5.5 system card says Anthropic now follows the Surge leaderboard in using Gemini 3.5 Flash as the grader, and that "previous system cards showed slightly lower scores due to using Claude Sonnet 4.6 as a judge only for Claude models."[8] Consistent with this, Fable 5.1 appears at 42.6% / 86.2% in its own system card and at 44.8% / 88.4% three weeks later, and Opus 5 moves from 29.6% / 83.0% to 29.8% / 83.4%.[6][8]
- Runs per task. Anthropic averages five runs, Surge's paper uses 20 trials per task, and Surge's README specifies ten runs per task for the leaderboard configuration.[1][4][5]
- Who ran it. Developers' internal runs do not always match Surge's. For example, Surge lists Claude Opus 5.5 at 66.3% and Claude Opus 5 at 27.3%, against Anthropic's 64.4% and 29.6-29.8% no-tools figures.[3][5][8]
- Tool harness. With-tools numbers depend on the scaffold. Surge's leaderboard lists DeepSeek V4.1-Flash at 11.6% under its tool-free protocol, while DeepSeek reports 78.9 with tools in a Claude Code harness.[3][12] DeepSeek's figure for Claude Opus 4.8 (65.0) also differs from Anthropic's own with-tools (75.0%) and no-tools (17.0%) figures for that model, and DeepSeek's chart does not say which setting it used.[5][11]
Tuesday Work Index
On 18 August 2026 Surge introduced the Tuesday Work Index, a composite "measure of professional intelligence" whose first version combines eight Surge benchmarks: Chartography (listed as "professional graphical reasoning"), HANDBOOK.md, Antidote, Hemingway-bench, ComplexConstraints, GDP.pdf, CoreCraft and Riemann-bench.[13] In its September 2026 write-up of Claude Fable 5.1, Muse Spark 1.3 and Gemini 3.8 Flash, Surge reported that Fable 5.1 took first place on Chartography at 46.2% (at that time) and that Muse Spark 1.3 fell from 32.1% to 27.6% relative to Muse Spark 1.2.[14]
Availability
The evaluation harness is public on GitHub (surge-ai/chartography), first committed on 16 July 2026. Its committed task pack contains only two dummy sample tasks for local dry runs; the README directs users to a Hugging Face dataset, surgeai/chartography, for the released set.[4] The paper states that Surge releases "all tasks, images, provenance metadata, and evaluation code."[1] When checked on 29 September 2026, however, the surgeai/chartography dataset URL returned a "page not found" error to anonymous visitors and was not among Surge AI's public Hugging Face datasets, and the GitHub repository did not declare a license.[4][15] The arXiv paper itself is published under CC BY 4.0.[1]
See also
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36 ^37 ^38 ^39 ^40 ^41 ^42 ^43 ^44 ^45 ^46 ^47 ^48Garre, Suhaas; Mutty, Chris; Mehta, Sushant; Chen, Edwin. "Chartography: A Benchmark for Professional Chart Understanding." arXiv:2608.10677, submitted August 11, 2026. arxiv.org/...2608.10677
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Surge AI. "Chartography: A Benchmark for Professional Chart Understanding." Surge AI blog, July 16, 2026 (page also dated September 16, 2026). surgehq.ai/...chartography
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Surge AI. "Chartography Benchmark | Surge AI" (leaderboard). Accessed September 29, 2026. surgehq.ai/...chartography
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14Surge AI. "surge-ai/chartography" (README). GitHub, first commit July 16, 2026. github.com/...chartography
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Anthropic. "System Card: Claude Opus 5." July 24, 2026. Section 8.12.1. www-cdn.anthropic.com/...s%205%20System%20Card.pdf
- ^1 ^2 ^3 ^4 ^5 ^6Anthropic. "System Card: Claude Fable 5.1 & Claude Mythos 5.1." September 1, 2026. Section 8.14.1. www-cdn.anthropic.com/...205.1%20System%20Card.pdf
- ^1 ^2Anthropic. "Introducing Claude Opus 5.5." September 22, 2026. anthropic.com/claude-opus-5-5
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Anthropic. "System Card: Claude Opus 5.5." September 22, 2026. Section 8.13.1. anthropic.com/claude-opus-5-5-system-card
- ^1 ^2 ^3 ^4 ^5Anthropic. "Introducing Claude Sonnet 5.5." September 28, 2026. anthropic.com/claude-sonnet-5-5
- ^1 ^2 ^3 ^4 ^5 ^6Anthropic. "System Card: Claude Sonnet 5.5." September 28, 2026. Section 8.13.1. anthropic.com/claude-sonnet-5-5-system-card
- ^1 ^2 ^3DeepSeek. "DeepSeek-V4-Flash-Vision-Exp Release: Multimodal API Now Live." DeepSeek API Docs, August 21, 2026. api-docs.deepseek.com/...news260821
- ^1 ^2 ^3 ^4DeepSeek-AI. "deepseek-ai/DeepSeek-V4.1-Flash" (model card). Hugging Face, September 2026. huggingface.co/...DeepSeek-V4.1-Flash
- ^Surge AI. "Introducing the Tuesday Work Index: Can AI Get Through an Ordinary Day?" Surge AI blog, August 18, 2026 (page also dated September 24, 2026). surgehq.ai/...tuesday-frontier-work-index
- ^1 ^2Surge AI. "Fable 5.1, Muse Spark 1.3, and Gemini 3.8 Flash on the Tuesday Work Index." Surge AI blog, September 10, 2026 (page also dated September 16, 2026). surgehq.ai/...fable-5-1-tuesday-work-index
- ^Hugging Face. Public datasets of the surgeai account (API listing). huggingface.co/...datasets. Accessed September 29, 2026.
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,403 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: xg12 V4 independent verification 29 Sep 2026: all 12 domain rows, 30 launch rows, 16 leaderboard rows checked; 0 material, 6 minor fixed.
Cite this page: AI Wiki. "Chartography." aiwiki.ai, updated 29 Sept 2026, fact-checked 29 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/chartography