# Ox Alpha

> Source: https://aiwiki.ai/wiki/ox_alpha
> Updated: 2026-08-22
> Fact-checked: 2026-08-22
> Categories: AI Code Generation, AI Models, Large Language Models
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Ox Alpha." aiwiki.ai, 22 Aug 2026. https://aiwiki.ai/wiki/ox_alpha
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**Ox Alpha** is the preview name of an anonymous [large language model](https://aiwiki.ai/wiki/large_language_model) made available through [OpenRouter](https://aiwiki.ai/wiki/openrouter) and [OpenCode](https://aiwiki.ai/wiki/opencode) on Aug. 20, 2026. OpenRouter describes it as a reasoning model for coding and sustained agent work. The service also states that an unnamed third-party provider develops and operates the model, while OpenRouter only routes requests.[1][2]

The public listing documents a 1,048,576-token context window, text, image, and video input, text output, and a maximum output of 131,072 tokens on the OpenRouter route. OpenCode announced a one-week free preview with a separate route and privacy policy.[3][5] These are distribution and interface facts. They do not identify the developer or disclose the model's architecture, parameter count, training data, compute, weights, or license.

As of Aug. 22, 2026, neither distributor had named the developer or linked a developer-authenticated checkpoint, technical report, model card, or reproducible evaluation package. Ox Alpha therefore remained a temporary stealth preview rather than a documented model release.[2][6]

## Key facts

| Field | Detail |
| --- | --- |
| Public preview | Aug. 20, 2026[1][5] |
| Developer | Undisclosed third-party provider[2] |
| OpenRouter model ID | `stealth/ox-alpha`[3] |
| OpenCode Zen model ID | `x-preview-f-free`[6][7] |
| OpenCode Go model ID | `ox-alpha-free`[8] |
| Stated focus | Coding, sustained agentic work, and workflows combining text with visual context[1][2] |
| Context window | 1,048,576 tokens on the listed routes[3][5] |
| Maximum output | 131,072 tokens on OpenRouter[3][4] |
| Input and output | Text, image, and video input; text output[3] |
| OpenRouter reasoning control | Mandatory; max, high, and low effort settings[3] |
| OpenRouter and Zen token price | Free at the research cutoff[2][6] |
| OpenCode Go access | Go subscription ($5 for the first month, then $10 per month); Ox Alpha is listed without token charges for a limited time[8] |
| Public artifact status | No checkpoint, paper, license, or model card linked by the distributors or a disclosed developer[2][6] |

## Preview and distribution

### OpenRouter route

OpenRouter's announcement presented Ox Alpha as a stealth model for [AI code generation](https://aiwiki.ai/wiki/ai_code_generation) and long-running agent workflows. The launch post gave two main specifications: a one-million-token context window and text, image, and video input. A follow-up in the same thread said the route was free and that the provider did not train on prompts or completions.[1]

The current model page is more precise about the parties. It says the model is developed and operated by a third-party provider that chose to remain anonymous during the preview. OpenRouter says it is not the developer, owner, or provider.[2] The endpoints API lists one provider under the generic name `Stealth`; it does not disclose a company or laboratory.[4]

OpenRouter prices both prompt and completion tokens at zero for the preview. Its stealth terms make clear that this category of access is time-limited and may be removed with or without notice.[2][9] The zero price is therefore a dated preview condition, not a permanent price or a promise that the route will continue.

### OpenCode routes

OpenCode announced that Ox Alpha would be free for the next week. Its launch post repeated the one-million-token and multimodal-input claims and advertised generous rate limits, near-unlimited use, zero data retention, and capacity for 100 trillion tokens per day.[5]

That final figure is an operator capacity claim. It is not a count of tokens actually processed, a per-user allocation, a measured tokens-per-second rate, a service-level guarantee, or evidence about the model's parameter count. OpenRouter publishes changing traffic and performance counters for the route, but those short-window measurements are different from OpenCode's aggregate capacity statement and are not stable model specifications.[2][5]

OpenCode Zen listed Ox Alpha's input, output, and cached-read prices as free. OpenCode Go has a separate plan charge: $5 for the first month and $10 per month afterward. The Go catalog showed no token charges for Ox Alpha and labeled it free for a limited time.[6][8] The word "free" therefore described the model's preview token price, not the absence of every possible subscription charge.

OpenCode uses different identifiers in its two product catalogs. Its Zen documentation maps Ox Alpha Free to `x-preview-f-free` on an OpenAI-compatible chat-completions endpoint.[6][7] OpenCode Go maps the same preview name to `ox-alpha-free` on its own endpoint.[8] The `owned_by: opencode` value in OpenCode's model API identifies the API namespace. It does not say that OpenCode trained the underlying model.[7]

## Documented interface and limits

The OpenRouter models API describes a text, image, and video to text route. It lists a 1,048,576-token context length and a 131,072-token maximum completion. The endpoint accepts controls for maximum tokens, sampling temperature, top-p and top-k sampling, response format, tools, tool choice, and reasoning effort.[3][4]

Reasoning is mandatory on this route, enabled by default, and exposed at max, high, or low effort. The listed defaults are temperature 1 and top-p 0.95.[3] These fields describe how a client can call the hosted route. They do not show whether the underlying model is dense or sparse, how it was trained, how many parameters it has, or how its internal reasoning process works.

Tool and tool-choice parameters make the route usable in [AI agents](https://aiwiki.ai/wiki/ai_agents) that implement [function calling](https://aiwiki.ai/wiki/function_calling). Parameter support alone does not establish reliable tool selection, correct structured output, or success on long software-engineering tasks. Those behaviors require controlled evaluation with a specified agent, environment, and grader.[3][4]

The modality field has a similarly narrow meaning. The hosted route fits the interface definition of a [multimodal model](https://aiwiki.ai/wiki/multimodal_model): Ox Alpha accepts text, images, and video and returns text. The listing does not document audio input, audio output, image generation, or video generation.[3] It also does not state how a video is sampled, whether its audio track is processed, or how visual tokens count against the context budget.

## Data handling by route

Ox Alpha's access routes have different published data-handling statements. They should not be collapsed into one model-wide privacy claim.[2][5][6][8]

| Route | Retention statement | Training statement | Scope |
| --- | --- | --- | --- |
| OpenRouter | The anonymous provider retains prompts and completions | The route page says they are not used; the incorporated EULA authorizes training and improvement | OpenRouter route and Stealth Program terms[2][9] |
| OpenCode Zen | Zero retention | Not used for model training | OpenCode route[5][6] |
| OpenCode Go | 0 days | Not used for model training | OpenCode Go route[8] |

OpenRouter's route-specific notice does not amount to zero data retention, and the route page and incorporated EULA do not align cleanly on training. The Ox Alpha page says the provider retains prompts and completions but does not use them for training.[2] The EULA says Stealth Models collect user content for training and improvement, tells users who do not want that use to refrain from accessing them, and grants OpenRouter and the provider rights to use content for training, evaluation, and improvement. It also says user content may be collected and shared with the provider and that personal data in an input will be sent to it.[9] The published materials do not explain how the Ox Alpha-specific notice limits those broader terms. The supportable conclusion is nonzero retention and an unresolved training-policy conflict, not an unqualified no-training guarantee.

OpenCode's zero-retention statement applies to access through OpenCode. It cannot be used to describe an OpenRouter request, even when both routes reach the same preview name. The unknown developer also means that users cannot independently compare the provider's internal controls, audit reports, or jurisdiction with those of a named model vendor.[2][6][8]

## Evaluation evidence

No official Ox Alpha benchmark report accompanied the preview. Ben Davis posted an informal run of ten DeepSWE tasks and explicitly warned that such a small subset could have substantial variance.[10] The post did not provide the task IDs, random seed, benchmark version, agent configuration, model settings, complete trajectories, grader files, or full logs. Its resulting percentage cannot be treated as a model-wide DeepSWE score or a controlled head-to-head result.

DeepSWE contains 113 original, long-horizon software-engineering tasks across 91 active open-source repositories and five programming languages. Each task has a hand-written verifier intended to check requested behavior rather than one reference patch.[11] The public repository uses the Harbor task format and a separate verifier environment. It documents deterministic subset sampling and says its leaderboard runs use Pier with mini-swe-agent on Modal, with evaluation trajectories released for inspection.[12]

Those details matter because a coding-agent result measures a full system, not only model weights. The agent scaffold, tools, network access, time limit, token budget, sampling settings, retries, selected tasks, benchmark revision, and verifier execution can all affect the outcome. Comparisons produced under different configurations are not interchangeable. DeepSWE's relationship to [SWE-bench](https://aiwiki.ai/wiki/swe_bench) does not remove that requirement; it was designed in part to improve task originality and grading, not to make undocumented samples representative.[11]

At the cutoff, no full Ox Alpha DeepSWE trajectory package, official 113-task submission, or independent replication was located. The preview can be tested by users, but anecdotes and screenshots do not establish a benchmark rank.[10][12]

## Interpreting context and multimodal claims

A context-window number states how much tokenized input a route accepts. It does not show how reliably the model finds or combines information throughout that input. "Lost in the Middle" found that long-context language models often performed best when relevant information appeared near the beginning or end and worse when it appeared in the middle.[13] RULER extended long-context testing beyond simple retrieval to multi-hop tracing and aggregation, and reported substantial performance declines as context length and task complexity increased.[14]

Neither study tested Ox Alpha. They establish why the one-million-token listing should be described as capacity rather than proof of effective reasoning over one million tokens. A direct evaluation would need controlled tasks at multiple lengths and positions, with the exact route and settings recorded.[3][13][14]

Video input also needs task-level evaluation. LongVideoBench tests retrieval and reasoning over interleaved video and language inputs up to one hour, while [Video-MME](https://aiwiki.ai/wiki/video_mme) covers short, medium, and long videos across several content domains.[15][16] No Ox Alpha result on either benchmark was located. The supported-input field therefore establishes that a client may send video, not that the model understands long temporal sequences at a particular level.

## Identity and undisclosed details

OpenRouter labels the API tokenizer as `Other`, leaves the knowledge cutoff unset, and provides no Hugging Face repository ID. The endpoint API labels quantization as unknown.[3][4] None of these fields identifies a model family. Similar API parameters or [tokenization](https://aiwiki.ai/wiki/tokenization) behavior can occur across related serving stacks and do not establish authorship.

The preview also lacks public disclosures needed for conventional model documentation.[2][3][4][6]

| Undisclosed area | What is not established |
| --- | --- |
| Developer identity | Company, research laboratory, or training organization |
| Architecture | Dense or mixture-of-experts design, parameter counts, attention design, vision encoder, and tokenizer |
| Training | Data sources, compute, optimization, post-training, and knowledge cutoff |
| Artifacts | Weights, source repository, model card, technical report, and license |
| Evaluation | Official benchmark suite, raw outputs, safety testing, and independent replication |
| Operations | Hosting hardware, geographic processing, long-term price, and availability after the preview |

Community attempts to infer the developer from output wording, model self-identification, error strings, geopolitical answers, and tokenizer comparisons do not fill these gaps. A first-party announcement by the developer or an explicit relabeling by the distributors would be evidence. Social consensus around a guess would not.[2][3][4]

## Availability and reproducibility

OpenCode framed its initial offer as one week. OpenRouter's stealth terms say availability is limited and may end at the provider's request or OpenRouter's discretion.[5][9] A current API entry therefore records access at a point in time rather than a durable release schedule.

Without authenticated weights, a model card, exact training information, and a full evaluation configuration, outside researchers cannot reproduce the model or separate its behavior from the serving route. They can run black-box tests, but those tests remain dependent on a mutable endpoint whose provider and version are not disclosed. Any later identity reveal, permanent release, pricing change, or artifact publication would require a new factual review rather than an inference from the preview.[2][3][6][9]

## References

1. OpenRouter. "New stealth model: Ox Alpha." X, Aug. 20, 2026. https://x.com/OpenRouter/status/2090544970923184269
2. OpenRouter. "Ox Alpha: API Pricing & Providers." Accessed Aug. 22, 2026. https://openrouter.ai/stealth/ox-alpha
3. OpenRouter. "Models API." Ox Alpha entry, accessed Aug. 22, 2026. https://openrouter.ai/api/v1/models
4. OpenRouter. "Ox Alpha Endpoints API." Accessed Aug. 22, 2026. https://openrouter.ai/api/v1/models/stealth/ox-alpha/endpoints
5. OpenCode. "Ox Alpha (stealth model) is free for the next week." X, Aug. 20, 2026. https://x.com/opencode/status/2090544355824038300
6. OpenCode. "Zen." Accessed Aug. 22, 2026. https://opencode.ai/docs/zen/
7. OpenCode. "Zen Models API." Accessed Aug. 22, 2026. https://opencode.ai/zen/v1/models
8. OpenCode. "Go." Accessed Aug. 22, 2026. https://opencode.ai/docs/go/
9. OpenRouter. "Stealth Program End User License Agreement." Updated July 6, 2026. https://openrouter.ai/terms/stealth
10. Ben Davis. "I ran this thing through 10 tasks on DeepSWE." X, Aug. 21, 2026. https://x.com/davis7/status/2090655207831298095
11. Wenqi Huang, Charley Lee, Leonard Tng, and Serena Ge. "DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks." arXiv:2607.07946, July 8, 2026. https://arxiv.org/abs/2607.07946
12. Datacurve AI. "DeepSWE." GitHub repository, accessed Aug. 22, 2026. https://github.com/datacurve-ai/deep-swe
13. Nelson F. Liu et al. "Lost in the Middle: How Language Models Use Long Contexts." Transactions of the Association for Computational Linguistics, 2024. https://arxiv.org/abs/2307.03172
14. Cheng-Ping Hsieh et al. "RULER: What's the Real Context Size of Your Long-Context Language Models?" COLM 2024. https://arxiv.org/abs/2404.06654
15. Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. "LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding." arXiv:2407.15754, July 22, 2024. https://arxiv.org/abs/2407.15754
16. Chaoyou Fu et al. "Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis." arXiv:2405.21075, revised May 30, 2025. https://arxiv.org/abs/2405.21075
