Citation and evidence

Odyssey-3

16 min full readUpdated 13 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI ModelsEmbodied AIRoboticsWorld Models

Cite this article

Odyssey-3 is a foundation world model announced on September 15, 2026 by Odyssey, an AI lab founded in 2023 by Oliver Cameron and Jeff Hawke.[1][6] Odyssey describes it as an autoregressive diffusion transformer "trained to simulate highly diverse scenarios," and presents it as "an early example of a single learned intelligence operating across many different physical and virtual systems."[1] The announcement post, titled "Introducing Odyssey-3: A General-Purpose Physical Intelligence" and bylined to Cameron and Hawke, reports experiments in which the same pretrained model was used as the base for policies that control robot arms, humanoid robots, a road vehicle, a simulated drone, and video game characters, and in which the model itself serves as a generated environment for training other AI agents.[1]

The distinguishing claim is not a benchmark score but a training recipe. In the driving, drone and gaming experiments Odyssey says the pretrained world model stays frozen and only a comparatively small policy or "action decoder" is trained on paired observations and actions from the target system.[1] Odyssey argues this lets a new embodiment be adapted with "only a few hours of experiential data" because the broad knowledge of physics, dynamics and cause and effect already sits in the backbone.[1]

At announcement Odyssey said it was "excited to release it publicly in the coming weeks," so the model was not generally available on the day of the post.[1] Odyssey's research page links a technical report for its Starchild-1 model and an arXiv preprint for PROWL-1, but lists only a "Learn More" link pointing back to the announcement for Odyssey-3.[4] The newsletter AlphaSignal, reviewing the launch, wrote that at the time of its publication "no public download or pricing had been announced" and that Odyssey had "not specified the release format, license, checkpoint access, API limits, or hardware requirements."[3]

Announcement and framing

Odyssey published the Odyssey-3 post on its news index on September 15, 2026.[6] The company's X account posted a video of roughly 97 seconds the same day, writing: "Today we're unveiling Odyssey-3, a big step forward for foundation world models. It can control robots, power humanoids, drive cars (on the roads of India!), train AIs, pilot drones, and even play video games."[2] In a snapshot taken on September 16, 2026, the post showed roughly 1.58 million views, about 3,300 likes and around 220 replies.[2]

Cameron and Hawke open the post by tying Odyssey-3 to their own history in autonomous driving and robotics in the 2010s, and argue that the field instead produced "increasingly specialized systems, each trained for a relatively narrow domain and often requiring enormous amounts of task-specific data."[1] They position a world model as the shared substrate that avoids that specialization, and introduce the term "physical agents" for systems that use one to interface with physical and virtual machines.[1] Elsewhere the company markets the same idea as "one foundation world model" powering "policies for robots, humanoids, cars, drones, video games, and more," and lists robotics, self-driving, drones, defense, gaming, energy, cyber and AI training among its application areas.[5]

Method

Odyssey's description of how a pretrained world model becomes a controller has two parts.[1]

The first is experiential data: recordings that pair a physical system's observations with the actions taken while performing a task, which Odyssey describes as "providing examples of how its controls are used."[1] The second is an action decoder, which Odyssey defines as a learned output component attached to the world model that translates Odyssey-3's internal representations into the actions a given system requires.[1] Only the decoder, or a small policy on top of the model's representations, is described as being trained on the target system's data.

The driving section is the most explicit about what stays fixed: "The pretrained world model remains frozen throughout training, extracting visual representations that feed into a relatively small driving policy, which predicts waypoints ahead of the car."[1] The drone experiments are described as following "the same training recipe as our driving experiments," and the gaming policies are trained "keeping the pretrained world model frozen."[1] The post does not state explicitly whether the backbone is also frozen for the robot arm and humanoid work. AlphaSignal's summary generalizes the pattern to all five domains, describing a "frozen backbone" that "controls robot arms, humanoids, cars, drones, and video game characters via lightweight action decoders."[3]

Reported demonstrations

All of the figures below come from Odyssey's own announcement; the company did not publish success rates, trial counts or variance for any of them.[1][3]

DomainAdaptation data reportedReported result
Robot arms"tens of hours of robot demonstrations"Control of "a variety of robot arms" on multi-step tasks, plus recovery behaviors absent from the training demonstrations
Humanoids (with Flexion)"tens of hours of humanoid teleoperation data"Real-time task execution; better generalization to lighting changes than "the VLA baselines tested"
Driving"only 20 hours of simulated driving data"Closed-loop autonomous driving on streets in India; sim-trained policies traveled about 77% as far between safety-driver interventions as policies trained on real footage
Drones"tens of hours of simulated drone data"Stable flight and obstacle avoidance in a simulated indoor setting
Video gamesGameplay recordings paired with keyboard and mouse input; about two hours of GTA footage for one transfer experimentExtended play in GTA V; movement in Red Dead Redemption 2 and motorcycle riding in Sleeping Dogs without additional policy training

Expanded article table

Robot arms

Odyssey reports that with tens of hours of demonstrations, Odyssey-3 learned to control several different robot arms on tasks issued as natural-language instructions, with examples including "Pour the cereal into the bowl," "Put the wipes into the box," "Close the screwbox," "Pour coffee into the espresso cup" and "Clean the plate with the wipe."[1]

The more interesting observation is about failure recovery. Odyssey says it observed "recovery behaviors that are absent from the training demonstrations, including reorienting a gripper after a missed grasp and retrieving an object dropped in an unusual position or orientation," and argues these suggest the model's learned physical understanding helps a robot respond to situations it was never shown.[1] The company frames this as evidence for sample-efficient robotics, "where collecting demonstrations of every possible failure and recovery quickly becomes impractical."[1] This is a qualitative claim: no success-rate comparison against an imitation learning baseline is given.

Humanoids and Flexion

Odyssey announced a research collaboration with Flexion Robotics, the Zurich company building a hardware-agnostic autonomy stack for humanoid robots, describing it as "a leader in general-purpose robot intelligence with expertise in reinforcement learning and whole-body control."[1][7] Odyssey says Flexion "carried out substantial research and engineering" on top of Odyssey-3 as a base model to produce the humanoid control policies shown, and that with tens of hours of teleoperation data the resulting system performed tasks in real time.[1] Demonstrated instructions include "Open the blue container and take out the cardboard box" and "Move the plate to the center of the table and place the mug on top of it."[1]

The comparative claim is narrow and worth quoting exactly: "In our evaluations, these policies generalize better to environmental changes than the VLA baselines tested, continuing to execute tasks under lighting changes that cause baseline policies to fail."[1] Odyssey does not name the vision-language-action baselines, give their sizes, or report task-level scores.

Flexion co-founder and CEO Nikita Rudin is quoted in the post: "What excites us about Odyssey-3 is the opportunity to build on physical knowledge acquired far beyond a robot's own demonstrations. Combining that foundation with our research in humanoid learning and control opens up exciting possibilities for how quickly robots can acquire useful skills and adapt to unfamiliar situations."[1]

Driving

Odyssey says that with only 20 hours of simulated driving data, Odyssey-3 drove a car "in closed loop on the streets of India, generating driving trajectories in real time," with the frozen backbone feeding a small waypoint-predicting policy.[1] Example prompts include "Take the first roundabout exit," "Drive along the road" and "Do a 180 degree turn towards the left."[1]

The headline number is a relative one. Odyssey compared policies trained entirely in simulation against policies trained on real driving footage, evaluating both "on busy roads with frequent distractions," and reports that both "navigated bends while maintaining their lane, handled vehicles overtaking on either side, and turned at busy junctions."[1] It then states: "On real roads, the policies trained entirely in simulation traveled about 77% as far between safety-driver interventions as those trained on real footage."[1]

That figure is a ratio between two of Odyssey's own policies, not an absolute safety result. A safety driver was present and interventions were counted, and the post gives no absolute distance between interventions, no route description, no traffic conditions and no number of trials. AlphaSignal made the same point in its review, listing "absolute intervention-free distances, route details, traffic conditions, trial counts, and variance" as missing.[3]

Drones

Following the driving recipe, Odyssey trained an aerial navigation policy to produce flight waypoints from recent camera observations, the drone's motion state and a high-level navigation prompt, such as "Take off, navigate under the table, and land behind the bookcase."[1] With tens of hours of simulated drone data, the company says the policy "demonstrated stable flight while avoiding obstacles in a simulated indoor setting."[1]

Odyssey also ran a diagnostic with the backbone's weights frozen and no policy training, visualizing the model's predictions for aerial navigation tasks. In those qualitative rollouts it reports "plausible directional flight and motion around obstacles, alongside responses consistent with the world's semantic content," which it reads as evidence that pretraining supplies useful knowledge of spatial structure and motion.[1] The drone work as described is simulated throughout; AlphaSignal noted that physical flight was not reported.[3]

Video games

Game-playing policies were trained on gameplay recordings paired with keyboard and mouse inputs, again with the world model frozen: during play the policy observes recent frames, emits controls, and uses the resulting observations to choose the next action.[1] Odyssey reports extended sessions in Rockstar Games' GTA V, "alongside selected examples of driving, shooting, and hand-to-hand combat."[1]

The transfer result is the one most often repeated in coverage. Odyssey says GTA-trained policies "have produced movement in Rockstar Games' Red Dead Redemption 2 and motorcycle riding in Square Enix's Sleeping Dogs without additional policy training on either title," and describes one experiment in which "a mobility policy trained on approximately two hours of GTA footage produced horseback movement in Red Dead Redemption 2, applying controls learned in one game to a different character, vehicle, and environment."[1] AlphaSignal characterized this as zero-shot transfer but cautioned that a single example "cannot establish broad transfer, however, or distinguish abstract locomotion knowledge from transfer enabled by similar controls and visual patterns."[3]

Odyssey-3 as a training environment

The post also treats Odyssey-3 as a place for other AI systems to learn rather than only as a controller. Odyssey says the model "can generate environments that AIs can inhabit, taking actions and learning from their consequences," supporting what it calls a recursive learning system in which agents and the world model improve each other.[1]

This connects to PROWL-1, the framework Odyssey published on May 12, 2026, whose full name is Prioritized Regret-Driven Optimization for World Model Learning.[10] In PROWL an RL agent is rewarded for discovering failures in a world model's geometry, motion, visual consistency and action-conditioned dynamics, building an adversarial curriculum that the world model is then fine-tuned on; a Prioritized Adversarial Trajectory buffer re-ranks trajectories so that resolved failures are deprioritized.[10] The accompanying preprint, arXiv:2605.18803 by Ahmet H. Güzel, Jenny Seidenschwarz, Benjamin Graham, Jonathan Sadeghi, Jeffrey Hawke and Ilija Bogunovic, evaluates the method on a diffusion-based world model.[11] Odyssey's Odyssey-3 post frames the same loop as a safety argument as well, saying generated worlds "give us a way to study how increasingly capable agents behave when given the freedom to act."[1]

Partners

PartnerWhat Odyssey says it doesIndependent detail
Flexion (Flexion Robotics AG)"a leader in general-purpose robot intelligence with expertise in reinforcement learning and whole-body control"; built the demonstrated humanoid control policies on Odyssey-3Swiss company founded in December 2024 by researchers from ETH Zurich and NVIDIA, building a hardware-agnostic humanoid autonomy stack (see Flexion Robotics)[7]
Poke & Wiggle"a leader in robot data, robot policy analysis, and large scale benchmarking"; jointly evaluating Odyssey-3 "across different bodies, viewpoints, and controls"Poke & Wiggle GmbH, registered in Munich (Amtsgericht Munchen HRB 301692), whose site says "We are obsessed with breaking & improving robot policies"[8][9]

Expanded article table

Odyssey frames the Poke & Wiggle work as an open question rather than a result, asking "how consistently these capabilities hold up across different robots and environments" and saying the joint evaluation is meant to show "where its knowledge transfers, where it breaks down, and how those findings can guide further training."[1] No results from that collaboration had been published at the time of the announcement.[3]

Place in Odyssey's model line

Odyssey has shipped a series of world models since 2024, and Odyssey-3 is the first one the company presents primarily as a control backbone rather than as an interactive simulator.[4][6]

ModelAnnouncedOdyssey's description
ExplorerDecember 18, 2024A purpose-built world model for film and gaming
Odyssey-1May 28, 2025A world model with navigation input and stability
Odyssey-2October 27, 2025A general world model with open-ended inputs (Odyssey's current index wording; the 2025 releases were announced as interactive video and renamed as world models in early 2026, see Odyssey-2)
Odyssey-2 MaxApril 21, 2026A general world model that "materially advances the state-of-the-art in physical accuracy"
PROWL-1May 12, 2026An RL-driven adversarial framework for improving world models
Starchild-1May 2026A real-time multimodal world model
Agora-1May 18, 2026A multi-agent world model shared by several participants in real time
CaliBenchAugust 10, 2026A benchmark testing whether world models reproduce the true distribution of physical outcomes
Odyssey-3September 15, 2026A foundation world model that powers robots, cars, drones, AI training and video games

Expanded article table

The most recent quantitative results Odyssey published for the line came with Odyssey-2 Max, which the company reported scored 58.52 on the physics sub-score of VBench 2 and 93.02 on the physics subset of PAI-Bench, against 49.67 and 91.67 for Odyssey-2 Pro at roughly one third the size.[12] The Odyssey-3 post contains no comparable table.[1]

The company's product cards describe Odyssey-3 as "Our most powerful foundation world model yet, materially advancing the state-of-the-art in physical accuracy of world models," reusing the phrasing previously applied to Odyssey-2 Max.[4][12] Odyssey raised a $310 million Series B at a $1.45 billion valuation in June 2026, led by Natural Capital with participation from Amazon, GV, AMD Ventures, EQT and IQT, and said Amazon Web Services would become its preferred cloud provider with model optimization work on AWS Trainium alongside Amazon's Annapurna Labs.[13]

Context

Odyssey-3 sits at the intersection of two lines of work that had largely stayed separate. One is interactive world models used mainly as generators and environments, such as Google DeepMind's Genie 3, NVIDIA's Cosmos 3 family for physical AI, and Wayve's GAIA-4 driving world model. The other is robot foundation models and vision-language-action policies such as Figure's Helix, models from Physical Intelligence, and Reward AI's OM-1, which map observations and instructions directly to actions. Odyssey's pitch is that the generator and the policy backbone should be the same network, with per-embodiment training reduced to a decoder.

No head-to-head comparison exists to place Odyssey-3 against any of these. Odyssey published no benchmark scores for the model, named no baselines, and the planned external evaluation with Poke & Wiggle had not reported results.[1][3] Odyssey's own closing framing is an estimate rather than a measurement: the post argues that frontier world models "remain sub-scale, roughly two orders of magnitude behind language models," a figure the company gives without a citation.[1]

What was not published

The announcement is a demonstration post, and several things a reader would need to reproduce or rank the work are absent from it.[1][3]

MissingNotes
Technical report or paperOdyssey's research page links a PDF technical report for Starchild-1 and an arXiv paper for PROWL-1, but no document for Odyssey-3[4]
Model size and architecture detailThe post names the family (autoregressive diffusion transformer) but no parameter count, layer configuration or context length[1]
Training dataDescribed only as "a vast collection of visual observations of the world"; no composition, scale or licensing detail[1]
Quantitative resultsNo success rates, trial counts, variance, latency or control frequency for any of the five domains[1][3]
Baseline identitiesThe humanoid comparison is against unnamed "VLA baselines tested"[1]
Weights, license and pricingRelease promised "in the coming weeks"; no release format, license, checkpoint access or API limits stated[1][3]
Independent evaluationAll results are Odyssey's own; the Poke & Wiggle benchmarking had not reported[1][3]

Expanded article table

Two further caveats follow from the post's own wording. The driving result involved a safety driver whose interventions were the unit of measurement, so it is not a claim of unsupervised operation on public roads.[1] And the drone result is described entirely in simulation, so it is not evidence of sim-to-real transfer on flying hardware.[1][3]

Reception

Independent coverage in the first day after the announcement was thin. AlphaSignal published the most substantial outside write-up, summarizing the five demonstrations in a table and devoting a section headed "The evidence still has gaps," covering independent validation, driving metrics, physical coverage, collection costs, implementation details and generalization.[3] It also reported that Odyssey says organizations in robotics, autonomous driving, gaming and defense are already using the model, while noting the company "has not identified those users or described their deployments."[3] Odyssey's own site lists defense, energy and cyber among its application areas but does not name customers.[5]

AlphaSignal's assessment of what the launch would need to prove is the clearest external framing available: "Independent benchmarks will need to measure data efficiency, control reliability, latency, and sim-to-real transfer against body-specific baselines."[3]

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36 ^37 ^38 ^39 ^40 ^41 ^42 ^43 ^44 ^45 ^46 ^47Oliver Cameron and Jeff Hawke, "Introducing Odyssey-3: A General-Purpose Physical Intelligence," Odyssey, September 15, 2026. odyssey.systems/introducing-odyssey-3
  2. ^1 ^2Odyssey (@odysseyml), launch post, X, September 15, 2026. x.com/...2099900067356586276
  3. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16"Odyssey Builds One AI Backbone to Control Robots, Cars, Drones, and Games," AlphaSignal Newsroom, September 2026. alphasignal.ai/...rol-robots-cars-drones-and-games
  4. ^1 ^2 ^3 ^4"Leading World Model Research," Odyssey research index. odyssey.systems/research
  5. ^1 ^2Odyssey homepage. odyssey.systems
  6. ^1 ^2 ^3"The Latest from Odyssey," Odyssey news index. odyssey.systems/writing
  7. ^1 ^2"About," Flexion Robotics. flexion.ai/about
  8. ^Poke & Wiggle. pokeandwiggle.com
  9. ^"Imprint," Poke & Wiggle. pokeandwiggle.com/imprint
  10. ^1 ^2Jeff Hawke, "Introducing PROWL-1: Learning Through Discovery," Odyssey, May 12, 2026. odyssey.systems/introducing-prowl-1
  11. ^Ahmet H. Güzel, Jenny Seidenschwarz, Benjamin Graham, Jonathan Sadeghi, Jeffrey Hawke, Ilija Bogunovic, "PROWL: Prioritized Regret-Driven Optimization for World Model Learning," arXiv:2605.18803, May 11, 2026. arxiv.org/...2605.18803
  12. ^1 ^2Oliver Cameron, "Introducing Odyssey-2 Max: Scaled World Simulation," Odyssey, April 21, 2026. odyssey.systems/introducing-odyssey-2-max
  13. ^Oliver Cameron, "Our $310 Million Fundraise to Accelerate World Simulation," Odyssey, June 17, 2026. odyssey.systems/our-series-b

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

1 revision · v2 · 3,258 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Adversarial verification 2026-09-16 (cluster V1): every quote verified verbatim against the announcement; Poke & Wiggle registry details confirmed; 3 minor defects fixed

Cite this page: AI Wiki. "Odyssey-3." aiwiki.ai, updated 16 Sept 2026, fact-checked 16 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/odyssey_3

Suggest edit