Citation and evidence

TypeSafe AI

16 min full readUpdated 22 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI CompaniesAI ResearchDeveloper Tools

Cite this article

TypeSafe AI (styled TypeSafe) is a San Francisco AI lab founded in 2024 that builds what it calls "machine-native, composable AI": models whose outputs are typed decisions with probabilities for other software to consume, rather than text for people to read. The company came out of stealth on 15 September 2026 with a seed round of approximately $40 million led by DCVC and released its first model, Jev, into early access the same day [1][2]. Its co-founder and chief executive is Diogo Almeida, a former OpenAI researcher who is the fourth listed author of the InstructGPT paper; the company describes him as a co-inventor of RLHF and ChatGPT [1][15]. The other founders are Erik Gafni (CTO) and Sasha Sheng (COO) [1][6].

TypeSafe's position, set out in a company manifesto under the tagline "Build Prod, Not God", is that today's models are already intelligent enough to create large economic value but are "hard to build on" because they were trained to please people, and that a different training target, calibrated decisions rather than human preference, is what production automation needs [5]. Its multipliers for Jev's speed and cost (ranging from "up to 100 times" in the press release to "193.6x faster, 444.6x cheaper" on its home page) are the company's own measurements on its own workflow evaluations, and the company's evaluation site shows Jev behind several frontier models on accuracy while far ahead on cost and latency [1][7][8].

History

The company was founded in 2024 and worked in stealth for roughly two years before its public launch; Almeida's launch post says "after two years in stealth" and his announcement tweet says he "spent the last 2 years in stealth building a new way to train models (RLCD)" [4][16]. DCVC's portfolio page lists its first investment in TypeSafe as 2025 [3]. The company's GitHub organization, typesafe-ai, has a repository dating from May 2024, consistent with the founding year in the press release [17].

On 15 September 2026 the company announced the seed round and Jev through a Business Wire release, a launch post by Almeida, a Q&A published by DCVC general partner James Hardiman, and a tweet from Almeida that had drawn about 18.8 million views and 53,600 likes by the following day [1][2][4][16]. Forbes covered the launch under the headline "This $200 Million Startup Wants To Fix AI's Overconfidence Problem"; DCVC linked to that piece as its coverage of the round [2][19]. The Register ran a story the next morning under the headline "TypeSafe AI debuts model for machines that plays Doom" [10].

DateEventSource
2024Company founded in San FranciscoPress release [1]
May 2024First repository created in the typesafe-ai GitHub organizationGitHub [17]
2025DCVC's first investment in TypeSafeDCVC portfolio page [3]
31 Mar 2026"Founders You Should Know" video interview with Almeida posted to the company blogTypeSafe blog [14]
19 Jun 2026Almeida's AI Council talk "AI: too good to be true, too bad to be useful" posted to the company blogTypeSafe blog, AI Council [13]
8 Aug 2026system-one-adapter-python repository created (MIT)GitHub [17]
4 Sep 2026typesafe-sdk-python and typesafe-sdk-js repositories created (MIT)GitHub [17]
10-11 Sep 2026"The Bitterest Lesson" and "Lies, Damned Lies, and Benchmarks" essays publishedTypeSafe blog [11][12]
15 Sep 2026Emerges from stealth; ~$40M seed led by DCVC; Jev enters early accessPress release, DCVC, launch post [1][2][4]

Expanded article table

Founders and team

PersonRoleBackground as described by TypeSafe
Diogo AlmeidaCo-founder and CEO"Co-invented RLHF and InstructGPT, the methods that lead to ChatGPT and GPT4"; previously at Google Brain [6]
Erik GafniCo-founder and CTORepeat founder (Ravel, "multi-modal AI for dna-sequencing"); early employee at Invitae and Freenome [6]
Sasha ShengCo-founder and COOFormer research engineer at Meta AI's FAIR, working on News Feed, AI Experiences and AI Research; published at NeurIPS and ECCV [6]
Ke DengChief of StaffPress contact on the launch release [1]

Expanded article table

The "co-inventor" wording for Almeida is the company's and his own; the press release calls him "co-inventor of RLHF/ChatGPT", the team page says he "co-invented RLHF and InstructGPT", and DCVC's post says he "co-invented some of the core techniques and products that launched the current boom: InstructGPT, ChatGPT, and GPT-4" [1][2][6]. What is independently verifiable is that Almeida is the fourth author of "Training language models to follow instructions with human feedback" (Ouyang et al., 2022), the InstructGPT paper [15]. In his launch post Almeida writes that at OpenAI he "helped build the methods that made language models useful at following instructions and talking with people" and that "that work ended up as the research behind ChatGPT" [4].

The team page says the company is "a close-knit, flat team from OpenAI, Google Brain, Meta/FAIR, Stripe, Airbnb, Plaid, Docker, and more" that works in person five days a week from an office "near the Embarcadero station" in San Francisco [6]. Its LinkedIn page lists the company size as 11-50 employees [20]. As of 16 September 2026 its Ashby job board carried five full-time San Francisco openings: Member of Technical Staff (Model Capabilities), Member of Technical Staff (Backend/Platform), Member of Technical Staff (Infrastructure, Kubernetes), Member of Staff, and Founding Marketer [18]. The job listings describe the stack as primarily Python, with TypeScript, Next.js and Tailwind CSS for front ends and Kubernetes for orchestration [18].

Funding

RoundDate announcedAmountLeadOther named investors
Seed ("Series Seed" in DCVC's wording)15 Sep 2026"$40 million" (headline); "approximately $40 million" (boilerplate)DCVC (James Hardiman, General Partner)None named in the press release, DCVC's post, or FinSMEs' summary [1][2][21]

Expanded article table

DCVC's portfolio page dates its first investment in TypeSafe to 2025, which indicates the firm's involvement predates the September 2026 announcement; neither DCVC nor the company says when the $40 million round closed [3]. Hardiman's quoted rationale is that TypeSafe is "turning increasingly capable models into technology that developers can reliably build into products at scale"; in his own post he frames the bet around "no-humans-in-the-loop automation" that "just hasn't shown up yet" [1][2]. Neither the release nor DCVC's post names other participants or a valuation; the only public valuation figure is the "$200 Million Startup" of Forbes' headline, and the team page says only that the company is "backed by top-tier investors" [1][2][6][19].

Thesis: machine-native, composable AI

TypeSafe's manifesto, titled "Composable AI: Build Prod, Not God", states a mission "to pave the shortest path to an AI-based economic revolution by making intelligence composable to catalyze a Cambrian explosion of intelligent software" [5]. Its argument runs as follows. Current models "have long since crossed the threshold of intelligence needed for creating massive economic value", yet "most software still isn't meaningfully intelligent", so "the bottleneck isn't raw intelligence. It's that today's intelligence is hard to build on" [5]. The manifesto compares chat-style AI to early "horseless carriages" that kept the shape of the thing they replaced, and argues that because RLHF "directly optimizes for human preference", the foreseeable consequence is "AI that requires humans in the loop instead of running in the background" [5]. The alternative it proposes is AI "as a primitive that any programmer can invoke for semantic judgement and decisions, while still using code for what it's best at: exact computation", which the manifesto's own footnote describes as the neuro-symbolic dream "cheekily summarized as 'smart if-statements'" [5].

The document lays out three steps: ship "the shape of machine-native composable AI with the highest possible intelligence-per-dollar"; make that AI "reliable enough to transform the economy via real automation"; then provide "higher-level intelligence abstractions that are stable enough to compose and layer upon" [5]. Its yardstick for an economic revolution is global total factor productivity growth "reaching 3% within five years and holding at that level for ten", and it says it is "not racing towards the ever-moving goalposts of 'AGI'" [5].

In DCVC's Q&A Almeida adds that intelligence should become "a utility, a basic building block of automation" and predicts that once AI has automated most of the economy "over 99% of AI calls will be made by and for software, not humans" [2].

Technology: System One Models and RLCD

TypeSafe calls its model class "System One Models", after the fast, intuitive System 1 thinking in Daniel Kahneman's Thinking, Fast and Slow [4]. According to the launch post, the company built "a new stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD)" [4]. An arXiv search for the phrase "Reinforcement Learning for Calibrated Decisions" returned no paper as of 16 September 2026, and the company's public description of RLCD is limited to the launch post and the home page, which lists "a new architecture, a new sampler, and a new training algorithm" [4][7]. The launch post's FAQ answers the question "Why was a new training algorithm needed?" only at the level of objectives: RLHF optimizes for "the text that a human rater prefers", which the company calls "the wrong task for automation", and RLVR "is great for tasks with simple programmatic verification, but most real-world judgement tasks don't fit into that shape" [4].

The developer-facing interface is documented at docs.typesafe.ai. A call sends a "state" (a string or JSON object) plus a set of typed "questions", and the model returns "typed values and probability distributions that your code can branch on, sort by, and route with" [9]. There are three question primitives [8][9]:

PrimitiveQuestion typeReturns
NoulYes or noA probability of "yes"
ChoiceOne option from a defined setA distribution over the options, plus a confidence value
ScoreA level on a scaleA score, a distribution over levels, plus a confidence value

Expanded article table

The launch post contrasts the approach with existing LLMs in a table whose main claims are the company's own: outputs are "type-safe structured values" (structured output whose shape is "defined in advance"), so "the model never makes type errors"; sampling is parallel, generating "all outputs in a single query" rather than one token at a time; and every answer carries "calibrated probabilities and confidence scores" [4]. The company states that the "no type errors" claim is guaranteed by schema matching and is "not empirical", which is also why it plots Jev's hallucination rate as zero [4]. The Register noted that this "really isn't a fair comparison as its output is not natural language" and "does not preclude the possibility of being incorrect" [10]. The company's home-page FAQ poses the question "Can Jev still get things wrong?" and its launch post asks early-access developers to report "where Jev works, and where it falls short" [4][7]. The launch post also reports that Jev supports a cardinality of up to 255 options per choice, with a two-stage scoring-then-choosing procedure for larger sets [4].

Jev

Jev, named for the economist William Stanley Jevons and the Jevons paradox, is TypeSafe's first public System One Model, released in waitlisted early access on 15 September 2026 [1][4]. The company's published price is $0.042 per million input tokens ("$42 per billion tokens") with output tokens free, and it quotes end-to-end response times of 70 to 500 milliseconds [4]. The press release says the model "delivers frontier-level intelligence at less than 100 milliseconds of latency and is up to 100 times faster and less expensive than other frontier models" [1]. TypeSafe concedes that it "can't prove it isn't subsidized" and says it will "need the long-term to prove the sustainability of our pricing" [4].

The speed and cost multipliers the company has published vary by venue and by comparison, and all are the company's own numbers:

FigureWhere it appearsWhat it compares
"Up to 100 times faster and less expensive"Press release [1]Other frontier models, unspecified
"20-200x faster", "40-400x cheaper (w/ output tokens free)"Almeida's launch tweet [16]Unspecified
"40x-200x faster"Launch post [4]Frontier models "for System One shaped queries"
"193.6x Faster, 444.6x Cheaper"Home page [7]Company workflow evals; launch post calls these "the higher end of real world gains" [4]
0.114 s vs 8.566 s; $0.000081 vs $0.013880Home-page side-by-side demo [7]GPT-5.6 Terra at default reasoning, on a short, simplified query the company says paints its model "in an advantageous light" [4]
"238x lower input price than Claude Fable 5.1"Home page [7]Input price only, against Claude Fable 5.1

Expanded article table

The company's evaluation site, evals.typesafe.ai, publishes four "workflow evals" (security incidents, agent-trace observability, invoice processing, customer service) in which every model runs the same decomposed decision workflow, and reference labels are the average of GPT-6 Astra and Claude Fable 5.1 at high thinking [8]. On the site's overview chart, averaged across the four workflows, Jev scores 67.8% agreement with those labels at $0.0004 and 0.4 seconds per case, while GPT-5.6 Sol scores 74.1% at $0.0836 and 23.3 seconds and Claude Opus 5 scores 73.1% at $0.1761 and 37.8 seconds; Claude Sonnet 5 matches Jev's 67.8% at $0.1174 and 78.1 seconds [8]. In other words the company's own data show Jev on the cost and latency frontier rather than at the top of the accuracy column. The launch post lists the caveats: the workflows were written by the company's own model capabilities team, the reference answers are biased toward OpenAI and Anthropic models, and the LLM comparisons run through TypeSafe's own "System One LLM" wrapper [4].

The company published two demonstrations alongside the launch: a bot that plays Doom from a structured text representation of game state (not images), making about ten queries a second at a cost the company puts at roughly $7 per hour, and a Wikiracing agent that chooses among "hundreds to thousands of links" per step [4][10]. The Register's headline led with the Doom demonstration [10].

Open-source tooling and documentation

TypeSafe publishes client libraries and helpers under the MIT license on GitHub: typesafe-sdk-python and typesafe-sdk-js (both created 4 September 2026), system-one-adapter-python (created 8 August 2026), which the README describes as "a drop-in replacement for typesafe_sdk's system_one evaluation API, backed by LLM APIs instead of TypeSafe" for cost, speed and intelligence comparisons, and skills, agent skills for building with the System One API [17]. The adapter is the "System One LLM wrapper" through which the company runs competing models in its evals [4][17]. The documentation site covers the primitives, "patterns" such as speculative fan-out and confidence-gated routing, cookbooks, and an API reference, and a hosted console with a playground is linked from the launch post [4][9]. Access is through the hosted API and console with an API key; the press release says early access is waitlisted, and no repositories appear under a TypeSafe organization on Hugging Face as of 16 September 2026 [1][9].

Public positions on evaluation

Two unbylined pre-launch essays on the company blog, written in the first person by an author who describes working on InstructGPT at OpenAI, set out positions that shape how the company reports results. "The Bitterest Lesson" argues that Rich Sutton's bitter lesson is "the tip of an iceberg", and that in practice "doing the right task > data > compute > algorithms"; it cites the InstructGPT result that much smaller models trained on the right objective were preferred over GPT-3 [11]. "Lies, Damned Lies, and Benchmarks" (published at the URL slug "antibenchmaxxing") argues that any public eval a lab can optimize against "inevitably gets benchmaxxed", cites the Llama 4 Arena episode, Vending-Bench behaviour, and the successive revisions of the Artificial Analysis index after GPT-6 Astra's launch as examples, and commits TypeSafe to "no standard benchmark table in our model releases", to evals that are "dated snapshots and immediately retired once posted", and to publishing "the evidence that looks bad for us" [12]. The caveat blocks under each result in the launch post follow that policy [4].

Reception

Thomas Claburn of The Register, writing on 16 September 2026, summarized Jev as a model that "doesn't chat" but "produces typed probabilistic decisions", judged it "more likely to be used for sorting customer service problems and other business workflows" than for games, and flagged both the hallucination framing and the open question of whether demand for tokens will be as broad as demand for energy [10]. The Rundown AI noted that agreement with GPT-6 Astra and Claude Fable 5.1 "gives only a partial view of whether Jev makes correct decisions" and that "response times under load and sustained pricing remain open questions" [22]. FinSMEs recorded the round as a $40 million seed led by DCVC with no other investors listed [21].

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16TypeSafe AI Emerges From Stealth With $40M in Funding With New Model for Composable AI, Business Wire, 15 September 2026 (syndicated copies at Yahoo Finance and Morningstar).
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9James Hardiman, TypeSafe emerges from stealth with a new way of doing AI, DCVC News & Insights, 15 September 2026.
  3. ^1 ^2 ^3TypeSafe, DCVC portfolio page (First Investment: 2025), accessed 16 September 2026.
  4. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23Diogo Almeida, Introducing System One Models & Jev, TypeSafe AI blog, 15 September 2026.
  5. ^1 ^2 ^3 ^4 ^5 ^6 ^7Composable AI: Build Prod, Not God, TypeSafe AI manifesto, accessed 16 September 2026.
  6. ^1 ^2 ^3 ^4 ^5 ^6 ^7Our Team, TypeSafe AI, accessed 16 September 2026.
  7. ^1 ^2 ^3 ^4 ^5 ^6TypeSafe AI home page, accessed 16 September 2026.
  8. ^1 ^2 ^3 ^4Workflow evals, TypeSafe AI, accessed 16 September 2026.
  9. ^1 ^2 ^3 ^4Introduction, TypeSafe AI documentation, accessed 16 September 2026.
  10. ^1 ^2 ^3 ^4 ^5Thomas Claburn, TypeSafe AI debuts model for machines that plays Doom, The Register, 16 September 2026.
  11. ^1 ^2The Bitterest Lesson, TypeSafe AI blog, 10 September 2026.
  12. ^1 ^2Lies, Damned Lies, and Benchmarks, TypeSafe AI blog, 11 September 2026.
  13. ^AI: too good to be true, too bad to be useful, Diogo Almeida talk, AI Council; video embedded on the TypeSafe blog, 19 June 2026.
  14. ^Diogo Almeida - Founders You Should Know, TypeSafe AI blog (embedded YouTube interview from the Founders You Should Know channel), 31 March 2026.
  15. ^1 ^2Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida et al., Training language models to follow instructions with human feedback, arXiv:2203.02155, 2022.
  16. ^1 ^2 ^3Diogo Almeida (@CompleteSkeptic), launch tweet, X, 15 September 2026.
  17. ^1 ^2 ^3 ^4 ^5 ^6typesafe-ai GitHub organization: typesafe-sdk-python, typesafe-sdk-js, system-one-adapter-python, skills; repository metadata via the GitHub API, accessed 16 September 2026.
  18. ^1 ^2TypeSafe AI job board, Ashby, accessed 16 September 2026.
  19. ^1 ^2This $200 Million Startup Wants To Fix AI's Overconfidence Problem, Forbes (The Prompt), 15 September 2026 (headline; article body is paywalled).
  20. ^TypeSafe AI, LinkedIn company page, accessed 16 September 2026.
  21. ^1 ^2TypeSafe AI Raises $40M in Seed Funding, FinSMEs, September 2026.
  22. ^TypeSafe launches Jev for AI decisions inside software, The Rundown AI, September 2026.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

1 revision · v2 · 3,198 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: xg04 V3 independent verification 2026-09-16 vs Business Wire, DCVC, company pages; 5 minor fixes applied

Cite this page: AI Wiki. "TypeSafe AI." aiwiki.ai, updated 16 Sept 2026, fact-checked 16 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/typesafe_ai

Suggest edit