# Octo (Robot Policy)

> Source: https://aiwiki.ai/wiki/octo
> Updated: 2026-09-27
> Fact-checked: 2026-09-27
> Categories: Embodied AI, Open Source AI, Robotics
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Octo (Robot Policy)." aiwiki.ai, 27 Sept 2026. https://aiwiki.ai/wiki/octo
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

Octo is an open-source generalist robot policy for manipulation. It uses a transformer and a diffusion action decoder to turn camera observations and a task description into robot actions. The Octo Model Team trained it on approximately 800,000 robot trajectories from a 25-dataset mixture drawn from [Open X-Embodiment](https://aiwiki.ai/wiki/open_x_embodiment).[1]

The research paper appeared at Robotics: Science and Systems in 2024.[2] Octo is a robot-control model, distinct from the similarly named [OctoAI](https://aiwiki.ai/wiki/octoai) company.

## Released models

The published 1.5 checkpoints have the following specifications.[3][4] Their reported parameter counts exclude the additional pretrained T5-Base language encoder.[2]

| Checkpoint | Reported parameters | Observation history | Predicted action chunk |
|---|---:|---:|---|
| `rail-berkeley/octo-small-1.5` | 27 million | Up to 2 timesteps | 4 successive 7-dimensional actions |
| `rail-berkeley/octo-base-1.5` | 93 million | Up to 2 timesteps | 4 successive 7-dimensional actions |

Both model cards identify an MIT license. The repository's MIT license permits reuse and modification subject to its notice requirements and disclaims warranties. These release terms describe the software and model artifacts; they should not be treated as a license for every independently contributed training dataset.[3][4][5]

## How the policy works

### Observations and task descriptions

The checkpoint input specification includes a primary RGB camera at 256 by 256 pixels and a wrist RGB camera at 128 by 128 pixels. Tasks can contain goal images or language instructions. Image processing uses a small convolutional encoder followed by 16-by-16 patches. Language passes through the [T5](https://aiwiki.ai/wiki/t5) tokenizer and language encoder. The model cards permit subsets of the documented observation and task keys, within the trained history length.[3][4]

The implementation separates task tokenizers, observation tokenizers, and output heads. Its blockwise causal attention allows observations at a timestep to use observations from that timestep or earlier ones, together with task information. Camera tokens at the same timestep can exchange information. Learned readout tokens collect information for the action head without changing the observation-token computation. This separation makes it possible to change an input module or an output head while retaining compatible pretrained components.[6]

Missing data has an explicit representation. A timestep mask distinguishes real observations from history padding; modality masks distinguish an absent camera or task channel from an available one. At the start of an episode, the missing previous observation should be masked rather than treated as a genuine earlier frame. The supplied history wrapper constructs the timestep mask.[17]

### Diffusion action prediction

Octo's action head implements a conditional denoising process. During training, it adds noise to demonstrated actions and learns to predict that noise using an embedding from the transformer. During inference, it starts with a noise sample and repeatedly refines it into an action chunk. The transformer runs once for this prediction; the repeated denoising steps occur in a smaller network conditioned on its output. The official `DiffusionActionHead` uses a multilayer perceptron with residual connections.[7]

This approach builds on [Diffusion Policy](https://aiwiki.ai/wiki/diffusion_policy), which formulates visuomotor control as conditional action generation. A distribution can represent different valid ways of acting in the same observed situation, rather than requiring a single average action. Diffusion Policy also combines action-sequence prediction with receding-horizon control: the robot executes part of a predicted sequence, observes again, and replans.[8]

For the released Octo checkpoints, predicting four actions does not mean predicting an entire task. The runtime may execute all four before another model call, use only the first, or combine overlapping predictions through temporal ensembling. These execution choices are separate from the checkpoint's prediction horizon.[17]

## Training data and reproducibility

Open X-Embodiment is a collection of robot datasets with standardized formats. Its project overview describes more than one million trajectories spanning 22 robot embodiments, including platforms beyond single robot arms. Those collection-wide figures are not Octo's training specification: the Octo release uses a selected mixture rather than every Open X-Embodiment dataset.[9]

The repository names its documented pretraining mixture `oxe_magic_soup`. The corresponding configuration lists 25 datasets, including Bridge, Fractal, Kuka, BC-Z, Language Table and Furniture Bench. It assigns different sampling weights to the datasets. Other mixtures in the same source file are separate configurations; their presence does not establish that the published Octo checkpoints were trained with them.[10]

The Small model card reports Fractal, Kuka and Bridge at 17.0% of each batch apiece. These are sampling proportions, not claims that each source contains the same number of trajectories. The remaining examples come from datasets with smaller reported batch shares.[4]

The released training stack uses [JAX](https://aiwiki.ai/wiki/jax) and includes data-loading examples for JAX and PyTorch. The README documents pretraining on a TPU v4-128 pod, reporting eight hours for Small and fourteen hours for Base, and approximately 1.2 TB for its preprocessed dataset. These are the project's reproduction figures, not hardware-independent runtime or storage guarantees.[17]

## Adapting Octo to another robot

[Fine-tuning](https://aiwiki.ai/wiki/fine_tuning) changes the policy using demonstrations from the target setup. The official adaptation example uses a simulated ALOHA cube-handover dataset to demonstrate a change in both observations and actions. It removes the wrist-camera tokenizer, adds a proprioceptive-input tokenizer, and replaces the action head with an L1-loss head producing 50-step chunks of 14-dimensional actions. Compatible pretrained parameters are copied into the modified model; new components receive newly initialized parameters.[11]

This example illustrates architectural flexibility, not an additional capability of the unchanged seven-dimensional checkpoint. A different action vector requires a compatible head and training data. Changing the number in a robot controller does not itself train the policy to use that action space.[11]

The general fine-tuning configuration exposes `full`, `head_only`, and `head_mlp_only` modes. It separately selects image-conditioned, language-conditioned, or multimodal training. Dataset configuration specifies camera keys, the language field, normalization, and the standardization transform. Its default training schedule uses 50,000 steps, but the supplied minimal adaptation example uses a different schedule; the examples are not interchangeable experiment recipes.[12][11]

The paper's adaptation experiments used about 100 demonstrations per target domain, with fine-tuning runs under five hours on an NVIDIA A5000. These are reported experiment settings, not requirements for arbitrary robots.[2]

## Reported experimental results

The project evaluated nine real-robot setups at four institutions, separating direct use of the pretrained policy from fine-tuning. The following are the project-reported language-conditioned zero-shot success fractions.[1]

| Policy | WidowX | UR5 | RT-1 robot |
|---|---:|---:|---:|
| RT-1-X | 0.20 | 0.35 | 0.60 |
| RT-2-X | 0.50 | Not reported | 0.85 |
| Octo | 0.50 | 0.70 | 0.80 |

Zero-shot here uses pretraining-distribution tasks and setups without fine-tuning, but varies initial conditions. The WidowX RT-2-X value came from earlier work; RT-2-X's authors supplied the RT-1 robot comparison.[2]

The six fine-tuning setups covered baking, coffee preparation, insertion, pick-up, bimanual manipulation and the Berkeley Coke task. Their reported average success rates were 0.72 for Octo, 0.20 for training from scratch and 0.15 for the VC-1 baseline.[1] Table I states 20 trials per domain, but Appendix F reports 10 for the bimanual task. The scratch baseline used a ResNet encoder and transformer, not randomly initialized Octo. Results apply to this evaluation suite.[2]

## Using a checkpoint

The `OctoModel` API loads checkpoints, builds task inputs and samples action chunks. Its action output is normalized unless the caller supplies the relevant unnormalization statistics. Those statistics must match the target dataset and embodiment. A numerically valid output array is therefore not automatically an appropriate physical robot command. The model object stores its configuration, example input batch and dataset statistics alongside its parameters.[13]

| Integration concern | What the official interface provides |
|---|---|
| Checkpoint loading | `OctoModel.load_pretrained` accepts a local checkpoint or an `hf://` repository identifier. |
| Task construction | `create_tasks` prepares text instructions or goal images and their masks. |
| Action sampling | `sample_actions` returns an action horizon and action dimension, with optional unnormalization. |
| Input inspection | `get_pretty_spec` summarizes expected image inputs, task inputs and output shape. |

The evaluation environment adapter expects observation dictionaries returned by `reset` and `step`. Camera key names must agree with the training configuration. An omitted input receives padding; adding a new sensor to the dictionary does not give the pretrained model a learned tokenizer for it.[14][6]

The real-robot example is specific to a WidowX setup and its controller. It includes workspace bounds, camera configuration, a control-step duration, history handling and action statistics from Bridge data. It also distinguishes blocking control from the non-blocking collection procedure. These details need adaptation when transferring the example to another system; they are part of the robot integration, not properties that the model infers from a language instruction.[15]

## Version 1.5 and research limits

The 1.5 release repeats language task tokens at each timestep and augments training instructions with paraphrases. Its release notes also record fixes to an attention-mask offset and image-augmentation random seeds, and disable dropout in the diffusion head. Results from different checkpoints should retain their version labels rather than being combined into a single performance claim.[16]

The original study reports weak use of wrist-camera information and a gap between language and goal-image conditioning. It also finds reduced performance on unfamiliar behaviors. Its experiments concern single-arm and dual-arm manipulation, not a demonstrated general navigation or whole-body-control system.[2]

Later research has used Octo as an experimental policy. For example, the ICLR 2026 paper *Policy Contrastive Decoding for Robotic Foundation Models* tests a training-free decoding method with Octo, OpenVLA and pi_0. That is a separate research intervention, not an official Octo checkpoint release. Its inclusion of Octo provides evidence of subsequent research use without establishing that the original policy has acquired the other models' capabilities.[18]

## References

1. Octo Model Team. [Octo: An Open-Source Generalist Robot Policy](https://octo-models.github.io/). Project overview and reported evaluation results. Accessed September 27, 2026.
2. Ghosh, D., et al. [Octo: An Open-Source Generalist Robot Policy](https://arxiv.org/html/2405.12213v2). *Robotics: Science and Systems*, 2024; arXiv:2405.12213v2. [Conference record](https://www.roboticsproceedings.org/rss20/p090.html), DOI: 10.15607/RSS.2024.XX.090.
3. Robotic AI & Learning Lab, UC Berkeley. [Octo Base 1.5 model card](https://huggingface.co/rail-berkeley/octo-base-1.5). Hugging Face. Accessed September 27, 2026.
4. Robotic AI & Learning Lab, UC Berkeley. [Octo Small 1.5 model card](https://huggingface.co/rail-berkeley/octo-small-1.5). Hugging Face. Accessed September 27, 2026.
5. Robotic AI & Learning Lab Berkeley. [Octo MIT license](https://github.com/octo-models/octo/blob/main/LICENSE). GitHub, copyright 2023.
6. Octo contributors. [Octo transformer and module implementation](https://github.com/octo-models/octo/blob/main/octo/model/octo_module.py). GitHub. Accessed September 27, 2026.
7. Octo contributors. [Action-head implementation](https://github.com/octo-models/octo/blob/main/octo/model/components/action_heads.py). GitHub. Accessed September 27, 2026.
8. Chi, C., et al. [Diffusion Policy: Visuomotor Policy Learning via Action Diffusion](https://diffusion-policy.cs.columbia.edu/). Project accompanying the RSS 2023 paper and IJRR 2024 article.
9. Open X-Embodiment Collaboration. [Open X-Embodiment: Robotic Learning Datasets and RT-X Models](https://robotics-transformer-x.github.io/). Project and dataset overview. Accessed September 27, 2026.
10. Octo contributors. [Open X-Embodiment dataset mixtures](https://github.com/octo-models/octo/blob/main/octo/data/oxe/oxe_dataset_mixes.py). GitHub. Accessed September 27, 2026.
11. Octo contributors. [Fine-tuning with a new observation and action space](https://github.com/octo-models/octo/blob/main/examples/02_finetune_new_observation_action.py). Official code example. Accessed September 27, 2026.
12. Octo contributors. [Fine-tuning configuration](https://github.com/octo-models/octo/blob/main/scripts/configs/finetune_config.py). GitHub. Accessed September 27, 2026.
13. Octo contributors. [OctoModel API implementation](https://github.com/octo-models/octo/blob/main/octo/model/octo_model.py). GitHub. Accessed September 27, 2026.
14. Octo contributors. [Octo evaluation environments](https://github.com/octo-models/octo/blob/main/examples/envs/README.md). Integration documentation. Accessed September 27, 2026.
15. Octo contributors. [Evaluating a fine-tuned model on a robot](https://github.com/octo-models/octo/blob/main/examples/04_eval_finetuned_on_robot.py). WidowX example. Accessed September 27, 2026.
16. Octo contributors. [Version 1.5 release notes](https://github.com/octo-models/octo/releases/tag/v1.5). GitHub. Accessed September 27, 2026.
17. Octo contributors. [Octo repository README](https://github.com/octo-models/octo). Installation, pretraining, evaluation and FAQ. Accessed September 27, 2026.
18. Wu, S., et al. [Policy Contrastive Decoding for Robotic Foundation Models](https://arxiv.org/abs/2505.13255). ICLR 2026; arXiv:2505.13255v6, April 24, 2026.

