DesignArena
| Field | Value |
|---|---|
| Type | Crowdsourced benchmark and consumer platform for AI-generated design |
| Developer | Intelligence (San Francisco) |
| Founders | Grace Li (CEO), Kamryn Ohly (CTO) |
| Launched | July 2025 (public Show HN July 12, 2025) |
| Accelerator | Y Combinator, Summer 2025 batch |
| Ranking method | Bradley-Terry model over anonymous pairwise votes |
| Users | 5.3 million (TechCrunch, August 2026); company claims 5.5 million |
| Funding | $7.9 million seed led by Index Ventures (announced August 3, 2026) |
| Website | designarena.ai |
DesignArena (also written Design Arena) is a crowdsourced benchmark and consumer creation platform for AI-generated design, built by the San Francisco startup Intelligence. Users describe something they want made (a website, game, image, video, slide deck, or app), receive outputs from several anonymized AI models, and vote in head-to-head comparisons; the votes feed public leaderboards ranked with the Bradley-Terry statistical model.[1][2][15] The company describes the platform as "the first crowdsourced benchmark for AI-generated design" and says its community spans more than 190 countries.[3][4] On August 3, 2026, the team introduced Intelligence as the company behind the product and announced a $7.9 million seed round led by Index Ventures, alongside a self-reported claim of growing from $5 million to $60 million in annual recurring revenue in six months with a team of ten.[5][6]
History
DesignArena grew out of a failed first idea. Co-founder Grace Li has said the team, a group of college friends who started working together a few weeks before graduating in 2025, was originally building an AI game engine. The models produced functional games, but not fun ones, and the founders concluded there was no substitute for human judgment on look and feel. They built a "this-or-that" voting game to score generated outputs for themselves, found the tool more compelling than the game engine, and made the benchmark the product.[1][7]
The site first appeared publicly in a Show HN post on July 12, 2025, describing "a crowdsourced benchmark for AI-generated UI/UX."[8] The team went through Y Combinator's Summer 2025 batch as Design Arena; the official YC launch post on July 31, 2025 reported more than 47,000 users across 136 countries in the first four weeks, and framed the benchmark as the first step of a broader mission then described under the name Arcada Labs.[3][9] A follow-up Launch HN post on August 12, 2025 said coverage had grown from roughly 25 language models at launch to 54 language models, 12 image models, 4 video models, 22 audio models, and 22 vibe-coding tools, and noted the surprise finding that agentic coding tools not marketed as builders, such as Devin, were outperforming dedicated tools like Lovable, v0, and Bolt in the builder category.[7]
YC lists the founding team as Grace Li (CEO, Harvard computer science and neuroscience, previously at Apple) and Kamryn Ohly (CTO, Harvard computer science and education, previously at Apple).[3][9] In the Launch HN post the team described its background as "Apple and Nvidia."[7]
On August 3, 2026, Li announced the company identity Intelligence (intelligence.ai, @Intelligence_ai) together with the seed round. In the announcement tweet she claimed the company had scaled "from $5M to $60M ARR and 5.5M users across 190+ countries" in six months as a team of ten; these figures are the company's own and have not been independently audited.[5] TechCrunch, reporting the round the same day, put the user base at 5.3 million and attributed the $60 million ARR figure to Li.[1]
How it works
Tournament format
Each voting session on DesignArena runs as a small tournament. A user enters a prompt and a category (or one is inferred), and four models are sampled from the active pool, plus one backup. All four generate outputs from the identical prompt simultaneously. Two randomly paired outputs are shown side by side with model identities hidden; the user picks a winner, the second pair is judged the same way, and winner and loser brackets plus a tiebreaker produce a complete first-through-fourth ranking. Each session yields five pairwise votes that feed the leaderboards.[4]
Rankings are computed with the Bradley-Terry model, a statistical framework for pairwise comparison data. Model "strengths" are estimated iteratively (convergence threshold 0.0001, capped at 200 iterations), normalized, and converted to a displayed rating as 400 times the base-10 logarithm of strength. Every vote is weighted equally with no editorial adjustment; charts filter out models with fewer than 50 pairwise comparisons, and models are marked preliminary until they reach roughly 200 comparisons, varying by category. The site states that model identities stay anonymized during voting to prevent brand bias and that methodology and configurations are publicly documented.[4] Because the platform wants only human votes, it applies aggressive bot protection, including captchas, and outputs require a login, which also lets the company track how design preferences differ across regions and over time.[1][7]
Categories
As of August 2026 the platform runs leaderboards across code, image, video, audio, slide, and miscellaneous design tasks, including:[2][10]
| Group | Leaderboards |
|---|---|
| Web dev (non-agentic) | Overall, Website, UI Component, Game Dev, Data Visualization, 3D Design, Image to Website, Video to Website (single-file HTML outputs) |
| Web dev (agentic) | Full-Stack, Frontend, Image to Frontend (complete React applications with auth, databases, and backend tool use) |
| Game dev | Agentic HTML, Godot |
| Mobile dev | Android (Kotlin, run in an emulator), React Native |
| Media | Image, Image Editing, Graphic Design, Logo, Video, Video Editing, Image to Video, Text-to-Speech |
| Other | SVG, ASCII Art, Python-PPTX Slides, HTML Slides, Builders |
The agentic evaluations capture more than the final output: the harness logs agent traces, tool calls, user re-prompts, failures, and retries, while human voters compare the final rendered applications or playable games.[10]
Leaderboard snapshot
DesignArena's models page publishes per-model overall win rates. Figures below are from designarena.ai as of August 7, 2026 and change continuously:[2]
| Model | Developer | Overall win rate | Top arena |
|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 63% | Mobile Apps |
| Claude Sonnet 4.6 | Anthropic | 59% | Mobile Apps |
| Grok 4.5 | xAI | 57% | Android Native |
| Claude Opus 4.8 | Anthropic | 57% | Agentic Game Dev |
| GLM 5.2 | Zhipu AI | 56% | Fullstack |
| Gemini 3.5 Flash | 55% | SVG | |
| Kimi K2.6 | Moonshot AI | 53% | HTML Slides |
| Gemini 3.1 Pro Preview | 51% | SVG | |
| GPT-5.5 | OpenAI | 51% | Game Dev |
The media boards display an Elo rating instead of a raw win rate; the site describes it as starting at 1200 and calculated with the standard Bradley-Terry model.[17] On Design Arena's video leaderboard as of August 8, 2026, Gemini Omni Flash (Google) led 39 listed models with an Elo of 1385, ahead of FLUX 3 Video from Black Forest Labs at 1326, MiniMax H3 at 1318, and ByteDance's Seedance 2.0 Mini (1292) and Seedance 2.0 (1290).[17] Design Arena had publicized the FLUX 3 Video result in an August 6, 2026 post, reporting the model second overall on its Video Arena with an Elo of 1325, ahead of MiniMax H3 and behind Gemini Omni Flash, and calling it particularly strong at Image to Video, where the post ranked it fourth.[18] On the separate Image to Video board as of August 8, MiniMax H3 led at 1352, followed by xAI's Grok Imagine Video 1.5 Preview at 1324, with FLUX 3 Video and Seedance 2.0 tied at 1301.[17]
Company and business model
For consumers, DesignArena functions like a model router with a chat-style prompt box and format dropdowns: users get several candidate outputs and rank them, keeping the best one. TechCrunch describes the real value as sitting on the enterprise side, where AI labs pay for the stream of human preference data as feedback for their media-generating models. Li told TechCrunch that taste "was the missing bottleneck for a lot of these models to make improvements in the design space," and that the company closed its first major deal with a frontier lab about a week after realizing labs would pay for the data.[1] The Intelligence site invites labs to submit models for evaluation and says many AI labs have partnered with the company.[11] In its August 2025 Launch HN post the team had described the planned business as version testing as a service, quantifying improvements between product builds.[7]
Because voters are logged in, the company can segment preferences geographically and temporally; Li has noted, for example, that web dashboards in Asia tend toward a more maximalist design style.[1] The $60 million ARR figure and the user counts (5.5 million in the company's announcement, 5.3 million in TechCrunch's reporting) are company-provided.[1][5]
Funding
| Round | Date announced | Amount | Lead | Other participants |
|---|---|---|---|---|
| Seed | August 3, 2026 | $7.9 million | Index Ventures | Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others; the company's announcement also credits Y Combinator |
TechCrunch and FinSMEs both report the round as led by Index Ventures with participation from Conviction, A*, and Valkyrie; Li's announcement additionally named Y Combinator as a participant.[1][5][6] FinSMEs reports the capital is intended for model evaluation infrastructure, expanding the reviewer community, and data engineering pipelines for multimodal design validation.[6]
Competitive landscape
DesignArena belongs to a wave of arena-style, human-preference evaluation businesses. The closest analogue is LMArena, which grew out of the academic Chatbot Arena project and applies pairwise human voting primarily to text responses; it raised a $150 million Series A at a $1.7 billion valuation in January 2026, four months after launching its paid product.[1][12][13] Artificial Analysis publishes independent model benchmarks alongside its own arena-style human preference comparisons for media models.[16] Human-preference platforms are not guaranteed businesses: TechCrunch notes that Yupp, a comparable feedback platform that raised $33 million from a16z crypto and claimed over 1.3 million users, shut down in 2026 less than a year after launching.[1]
TechCrunch frames human evaluation services as a complement to automated benchmarks, which scale further but can be gamed or manipulated.[1]
Limitations and criticism
DesignArena's own methodology page calls the benchmark "a subjective framework": rankings measure aggregate community preference, not any objective standard of usability or correctness, and the company acknowledges its builder comparisons currently rely on one-shot prompts under controlled conditions, with multi-turn evaluation listed as a planned extension.[4] Models with few votes carry preliminary status, and low-volume models are filtered from headline charts.[4]
Arena-style leaderboards as a class have drawn academic criticism. "The Leaderboard Illusion" (April 2025) documented how private variant testing, selective score retraction, and unequal data access can distort Chatbot Arena rankings, an argument that applies pressure to any vote-driven leaderboard whose operator also sells access to participating labs.[14] DesignArena's dual role, selling preference data and evaluation services to the same frontier labs whose models it publicly ranks, mirrors the commercial tension that paper describes; no allegations of ranking distortion at DesignArena were identified as of August 7, 2026. The platform's headline traction metrics (ARR and user counts) likewise come from the company itself and vary between tellings, with the announcement tweet citing 5.5 million users and press coverage citing 5.3 million.[1][5]
See also
- Chatbot Arena
- LMArena
- Artificial Analysis
- WebDev Arena
- Model evaluation
- AI benchmarks
- Y Combinator
- Scale AI
References
- ^Russell Brandom, "Design Arena creators raise $7.9 million to bring taste to AI models", TechCrunch, August 3, 2026, techcrunch.com/...lion-to-bring-taste-to-ai-models
- ^"Models", Design Arena, accessed August 7, 2026, designarena.ai/models
- ^"Design Arena: World's largest crowdsourced benchmark for AI-generated design", Y Combinator company directory, accessed August 7, 2026, ycombinator.com/...design-arena
- ^"About / Methodology", Design Arena, accessed August 7, 2026, designarena.ai/about
- ^Grace Li (@grx_xce), announcement of Intelligence and the $7.9M seed round, X, August 3, 2026, x.com/...2084361692792934488
- ^"Intelligence Raises $7.9M in Seed Funding", FinSMEs, August 5, 2026, finsmes.com/...ligence-raises-7-9m-in-seed-funding
- ^"Launch HN: Design Arena (YC S25) - Head-to-head AI benchmark for aesthetics", Hacker News, August 12, 2025, news.ycombinator.com/item
- ^"Show HN: DesignArena - crowdsourced benchmark for AI-generated UI/UX", Hacker News, July 12, 2025, news.ycombinator.com/item
- ^"Design Arena - #1 Benchmark for AI Design", Y Combinator Launches, July 31, 2025, ycombinator.com/...arena-1-benchmark-for-ai-design
- ^"Leaderboards", Design Arena, accessed August 7, 2026, designarena.ai/leaderboard
- ^"Intelligence", company website, accessed August 7, 2026, intelligence.ai
- ^"LMArena lands $1.7B valuation four months after launching its product", TechCrunch, January 6, 2026, techcrunch.com/...nths-after-launching-its-product
- ^"LMArena Secures $150 Million Series A", Cooley LLP news coverage, January 6, 2026, cooley.com/...marena-secures-$150-million-series-a
- ^Shivalika Singh et al., "The Leaderboard Illusion", arXiv:2504.20879, April 2025, arxiv.org/...2504.20879
- ^"Design Arena", homepage, accessed August 7, 2026, designarena.ai
- ^"Artificial Analysis: AI Model & API Providers Analysis", Artificial Analysis (image, video, and image-editing Arena leaderboards), accessed August 7, 2026, artificialanalysis.ai
- ^"Video" and "Image to Video" leaderboards, Design Arena, accessed August 8, 2026, designarena.ai/...video
- ^Design Arena (@DesignArena), post on FLUX 3 Video's Video Arena ranking, X, August 6, 2026, x.com/...2085490371606557139
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 2,185 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Video leaderboard standings verified against Design Arena's own API on August 8, 2026.
Cite this page: AI Wiki. "DesignArena." aiwiki.ai, updated 7 Aug 2026, fact-checked 7 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/designarena