Reward AI
Reward AI is a robotics company that builds general-purpose manipulation policies meant to run on many different robot bodies [1][2]. The company gives its location as the Bay Area, California, on its X account [12]. The company was co-founded by Zipeng Fu, its chief executive, and Chen Wang, its chief technology officer, both of whom completed PhDs at Stanford University before starting it [3][4]. Its public tagline is "Frontier Intelligence for Any Robot" [1]. The company's first announced model, OM-1, was published on 14 September 2026 [5].
Reward AI's LinkedIn profile lists it as a privately held artificial intelligence company founded in 2025 with 11-50 employees [6]. The company has not published funding details, customer names or product pricing, and as of mid-September 2026 its website consisted of a home page, an about page and a single blog post [1][2][5].
Founders
Zipeng Fu describes himself as a co-founder and the CEO of Reward AI. He completed a computer science PhD at the Stanford AI Lab in two and a half years, advised by Chelsea Finn, and was a researcher at Google DeepMind working with Jie Tan. Before Stanford he took a master's in the Machine Learning Department at Carnegie Mellon and worked in its Robotics Institute with Deepak Pathak and Jitendra Malik, after a bachelor's in computer science and applied mathematics at UCLA advised by Song-Chun Zhu [3]. His publications include Mobile ALOHA, HumanPlus, Deep Whole-Body Control, Robot Parkour Learning and the Open X-Embodiment collaboration [3].
Chen Wang describes himself as a co-founder and CTO. He received his PhD from Stanford computer science working with Fei-Fei Li and Karen Liu, graduating in June 2025, after a bachelor's in computer science at Shanghai Jiao Tong University with Cewu Lu. He has also done research stints at Google DeepMind in 2024, NVIDIA Research in 2022 and MIT CSAIL in 2019 [4]. His papers include DexCap, Sequential Dexterity, MimicPlay, VoxPoser, ReKep and TRANSIC [4]. DexCap, a portable motion-capture system for dexterous manipulation presented at Robotics: Science and Systems in 2024, is the work Reward AI names as the basis for its Omnibody Hand capture device [4][5][7].
Stated approach
The about page, signed "Zipeng", frames the company around three problems, which it summarizes as getting a robot to earn its place "in any body, at full speed, alongside others" [2].
| Problem | What Reward AI says it is doing |
|---|---|
| One Model, One Data Interface, Any Body | Route all manipulation data through one common interface into one model that spans robot arms, legged humanoids and wheeled mobile manipulators, so hardware changes do not require retraining or new data collection |
| Human Efficiency, Then Beyond | Treat speed as a research problem rather than a later optimization, designing perception, prediction and high-frequency control together, with human efficiency as a milestone rather than a ceiling |
| Many Robots, One Shared Goal | Build mutual understanding between robots, starting with quadmanual manipulation in which four robots coordinate, and extending to mixed teams that hand off objects and divide roles without scripted choreography |
Two consequences follow from the first of those. Reward AI says its model's life is not split into pre-training and fine-tuning, and that no round of on-robot data collection stands between training and deployment; from this it argues that data collected today "appreciates instead of depreciates" because it will still train robot bodies that have not been designed yet [2]. The company also states that every video clip it publishes runs at 1x speed and is labeled as such [2].
OM-1
OM-1, which Reward AI glosses as Omnibody Model 1, was announced on 14 September 2026 through a blog post, a 2 minute 19 second video and a post on X [5][8][9]. The company describes it as a general-purpose robot policy trained on human demonstrations captured with a wearable seven-degree-of-freedom device called Omnibody Hand, with no teleoperation data and no on-robot experience used in training, and says the same policy runs on tabletop arms, industrial arms and humanoids [5][9].
The announcement contains no model weights, code, dataset, API, paper or benchmark results, and the only measurement reported in it characterizes the capture rig rather than the policy: Reward AI reports that augmenting visual-inertial hand tracking with electromagnetic sensing cut mean overshoot error by 60% at high speed, from 24.9 mm to 9.5 mm at the fastest of eight tested speeds [5][10]. Trade coverage on 14 and 15 September summarized the post and repeated its claims with attribution [10][11].
Team background
The about page carries an image reel headed "Selected past work by our team" without saying which people worked on which project. The thirteen images on that reel are filed under names corresponding to ALOHA, HumanPlus, DexCap, Sequential Dexterity, VIOLA, DenseTact, deep whole-body control, a cycloidal gearbox, ToddlerBot, a Google DeepMind robot, a Science Robotics cover and a Falcon 9 landing [2]. The company describes its staff as "engineers, company builders and researchers who pioneered robot learning in dexterous manipulation, mobile manipulation, and legged locomotion, spanning the full stack: from large foundation models to gearboxes, from hands to legs, from visual and tactile perception to high-frequency control" [6].
That background is the argument Reward AI makes for attempting a full-stack robot foundation model: the same team claims to cover the wearable capture hardware, the sensing, the policy and the low-level controller, which is a wider span than most companies in the field take on at once.
Public presence
| Channel | Detail |
|---|---|
| Website | rewardai.com [1] |
| X account | @RewardAI_, created 7 September 2026 [12] |
| Company page listing 11-50 employees, founded 2025 [6] | |
| YouTube | @reward-ai; the OM-1 film was posted there on 14 September 2026 [8] |
| Contact | general@rewardai.com, careers@rewardai.com [2] |
Reward AI's public record is thin. There is no third-party evaluation of its model, no disclosed investor, no named customer and no shipped product; the assessment of the company rests on the founders' published research record and on demonstrations the company filmed itself.
References
- ^1 ^2 ^3 ^4Reward AI, home page. rewardai.com
- ^1 ^2 ^3 ^4 ^5 ^6 ^7"Towards Human-Level Robot Intelligence and Beyond", Reward AI about page. rewardai.com/about
- ^1 ^2 ^3Zipeng Fu, personal website. zipengfu.github.io
- ^1 ^2 ^3 ^4Chen Wang, personal website. chenwangjeremy.net
- ^1 ^2 ^3 ^4 ^5 ^6Reward AI Team, "OM-1: Frontier Robot Intelligence, Learned Firsthand from Humans", Reward AI Blog, September 2026. rewardai.com/...OM-1
- ^1 ^2 ^3Reward AI, LinkedIn company page. linkedin.com/...rewardai
- ^Chen Wang, Haochen Shi, Weizhuo Wang, Ruohan Zhang, Li Fei-Fei, C. Karen Liu, "DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation", arXiv:2403.07788, Robotics: Science and Systems 2024. arxiv.org/...2403.07788
- ^1 ^2"OM-1: Frontier Robot Intelligence Learned Firsthand from Humans", Reward AI, YouTube, 14 September 2026. youtube.com/watch
- ^1 ^2Reward AI (@RewardAI_), announcement post, X, 14 September 2026. x.com/...2099553899804053992
- ^1 ^2Asif Razzaq, "Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data", MarkTechPost, 14 September 2026. marktechpost.com/...teleoperation-or-on-robot-data
- ^Alisa Davidson, "'One Model, One Data Interface, Any Body': Reward AI's OM-1 Learns Manipulation Straight From Humans", Metaverse Post, 15 September 2026. mpost.io/...arns-manipulation-straight-from-humans
- ^1 ^2Reward AI (@RewardAI_) profile, X. x.com/RewardAI_
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 1,197 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independently fact-checked against 29 cited and primary sources (285 claims). 14 defects found, 3 material, all corrected.
Cite this page: AI Wiki. "Reward AI." aiwiki.ai, updated 15 Sept 2026, fact-checked 15 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/reward_ai