Citation and evidence

NVIDIA OpenShell

33 min full readUpdated 65 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI AgentsAI SafetyDeveloper ToolsNVIDIAOpen Source AI

Cite this article

NVIDIA OpenShell is an open source runtime that executes autonomous AI agents inside policy-governed sandboxes, published by NVIDIA under the Apache License 2.0. It sits between an agent and the machine it runs on, deciding which files the agent may read or write, which network destinations it may reach, which binaries are allowed to make those calls, and where the credentials it uses may be sent. NVIDIA announced OpenShell at GTC 2026 in San Jose on 16 March 2026, alongside NVIDIA NemoClaw, a reference stack that runs always-on agents inside OpenShell sandboxes.[1][2][3] The project is written mostly in Rust. After six months of near-daily 0.0.x releases labelled alpha, NVIDIA published OpenShell 0.1.0 on 25 September 2026, removing the alpha badges and redesigning the sandbox boundary.[39][40] Three days later, on 28 September 2026, NVIDIA described OpenShell as "now broadly available" and made it the software half of the NVIDIA Open Agent Safety Platform, alongside the NVIDIA Sentry reference design for BlueField-4 data processing units.[36][37]

FieldValue
DeveloperNVIDIA
Repository created24 February 2026; announced at GTC 2026, San Jose, 16 March 2026
LicenseApache License 2.0
ImplementationRust; YAML policies compiled to Open Policy Agent (Rego) and evaluated for each outbound request[35]
Project status0.1.x stable release series since 25 September 2026; the 0.0.x series (March to August 2026) was labelled alpha, "single-player mode"
Latest tagged releasev0.1.2, 28 September 2026[60]
Host platformsLinux (x86_64, aarch64), macOS on Apple Silicon, Windows via WSL 2 (experimental)
Compute driversDocker, Podman, MicroVM, Kubernetes; Windows MXC listed as "coming soon" (September 2026)
Part ofNVIDIA Open Agent Safety Platform (announced 28 September 2026)
GitHub stars8,960, with 1,298 forks and 124 contributors (28 September 2026)

Expanded article table

The problem it addresses

NVIDIA's launch post argued that a stateless chatbot has no meaningful attack surface, while an agentic AI system with a persistent shell, live credentials, the ability to rewrite its own tooling, and hours of accumulated context is a different threat model. The post framed the failure mode directly: with full access and full autonomy, you have "a long-running process policing itself", with guardrails living inside the same process they are supposed to guard.[2]

That matters because the dominant defence in most agent products is textual. System prompts, refusal training and tool descriptions are all instructions to a model, and a manipulated model will ignore them. Prompt injection is the canonical case. Greshake and colleagues showed in 2023 that instructions planted in data an LLM-integrated application retrieves, not typed by the user, can steer the application; they set out a taxonomy of impacts including data theft, worming and information ecosystem contamination, demonstrated attacks against real systems including Bing's GPT-4-powered Chat, and observed that "processing retrieved prompts can act as arbitrary code execution".[24] Prompt injection, direct and indirect, is the top entry (LLM01) in the OWASP Top 10 for LLM Applications 2025, which says that given the stochastic influence at the heart of how models work "it is unclear if there are fool-proof methods of prevention" and lists mitigations including privilege control and least-privilege access; the same list ranks Excessive Agency (LLM06) separately.[25][55]

An agent's promise not to do something is therefore not a security boundary. If a poisoned tool description, a malicious README or a compromised skill can redirect the agent, then every credential that process holds is reachable, every subagent it spawns can inherit permissions it was never meant to have, and every third-party skill it installs is an unreviewed binary with filesystem access.[2] OpenShell's response is to move the decision point outside the agent process entirely. Ali Golshan, NVIDIA's senior director of AI software, who leads product efforts on OpenShell, put it this way to The New Stack: "if you want to give more and more autonomy to an agent, the lowest level of the stack should really be a sandbox... That agent should not be interacting directly with your operating system or host or network or infrastructure."[15][33]

Origins

Golshan cofounded Gretel, a synthetic data company NVIDIA was reported in March 2025 to have acquired for a nine-figure sum, above Gretel's most recent $320 million valuation; two of his coauthors on the OpenShell launch post, Alex Watson and John Myers, were also Gretel cofounders.[33][34][2] Myers, a senior director of software engineering, leads OpenShell engineering, and Watson helps lead product.[2] The New Stack reported in May 2026 that Golshan and his team had been building OpenShell for the past six months, putting the start of work in late 2025; the public repository was created on 24 February 2026, three weeks before the GTC keynote.[15][1] In September 2026 the authors wrote that they had been "building OpenShell over the past year".[37] NVIDIA positions OpenShell as part of the NVIDIA Agent Toolkit, a bundle that LangChain's launch announcement described as including Nemotron models, NVIDIA NeMo Agent Toolkit profiling and optimization, NIM microservices and NVIDIA Dynamo.[2][15][20]

Architecture

OpenShell has gone through two sandbox designs. The description in this section through "Audit trail" covers the 0.0.x series as documented up to the end of July 2026, except where it notes later changes; the redesign that shipped in 0.1.0 is described in The 0.1 redesign below.

In the 0.0.x series the repository documented four components.[1][15]

ComponentRole
GatewayControl-plane API that coordinates sandbox lifecycle, holds credentials and session state, and acts as the authentication boundary
SandboxThe isolated runtime itself, supervised in-workload and with all egress routed through a local policy proxy
Policy EngineEvaluates filesystem, network and process constraints, from the application layer down to kernel enforcement
Privacy RouterPrivacy-aware model routing that keeps sensitive context on sandbox-local compute (removed in 0.1.0)

Expanded article table

NVIDIA's March 2026 launch post named three components (the sandbox, the policy engine and the privacy router), while a 23 March NVIDIA Blog post by Golshan listed the sandbox, the policy enforcement engine and the gateway.[2][14] Inside each 0.0.x sandbox workload there were two trust levels. A supervisor started as root, prepared the isolation, ran the proxy, injected credentials and launched the agent; the agent child then ran as an unprivileged user with filesystem, process and network restrictions already applied. On Linux the supervisor cleared the capability bounding set during the privilege drop so later exec calls could not regain container-granted capabilities, and aborted the spawn if that set did not end up empty.[6]

How the isolation works

OpenShell layers several mechanisms rather than relying on one primitive.[6][5]

  • Filesystem. Landlock, the Linux security module for unprivileged path-based sandboxing, restricts which paths the agent can read or write. A compatibility setting chooses between best_effort (skip inaccessible paths, apply the rest) and hard_requirement (fail startup if any path or the required kernel ABI is missing).[7]
  • Process. The agent runs as a non-root user and group; root identities are rejected outright. Seccomp filters block dangerous syscalls, including raw socket paths that would bypass the proxy.[7][6]
  • Network. In the 0.0.x design, a network namespace forced ordinary egress through a local HTTP CONNECT proxy, which identified the calling binary through a /proc/net/tcp inode lookup plus /proc/{pid}/exe, walked parent process identifiers, and enforced SHA-256 binary integrity on a trust-on-first-use basis.[11][6]

Filesystem and process policy are static, applied before the agent's first instruction runs, and changing them requires recreating the sandbox. Network policy is dynamic and hot-reloadable on a running sandbox.[1][7][35]

Network evaluation is deny-by-default and ordered: identify the calling binary, reject hard-blocked destinations including unsafe internal IP ranges, match destination and binary against policy blocks, apply optional layer-7 rules, then allow, deny, audit or log. Explicit denies and hardening checks beat allow rules, and if nothing matches the request is denied.[5] For endpoints marked protocol: rest the proxy terminates TLS with the sandbox's own ephemeral certificate authority so it can inspect method and path; it can also evaluate WebSocket upgrades, GraphQL operations, and Model Context Protocol method, tool and parameter rules on request bodies.[6][5]

The New Stack's May 2026 article described the enforcement layer as using "Linux kernel primitives such as seccomp, eBPF, and Landlock".[15] OpenShell's own architecture documents name Landlock, seccomp, network namespaces and the userspace proxy, and do not describe an eBPF component, so that part is best read as a press summary rather than documented implementation.[5][6]

The policy file

Policies are declarative YAML that teams are meant to version-control and review like any other security control.[4] A short policy in the documented format looks like this:[7]

version: 1
filesystem_policy:
  read_only: [/usr, /lib, /etc]
  read_write: [/sandbox, /tmp]
landlock:
  compatibility: best_effort
process:
  run_as_user: sandbox
  run_as_group: sandbox
network_policies:
  my_api:
    name: my-api
    endpoints:
      - host: api.example.com
        port: 443
        protocol: rest
        enforcement: enforce
        access: full
    binaries:
      - path: /usr/bin/curl

The binding of endpoints to binaries is the distinctive part: a connection is allowed only when the destination and the calling executable both match entries in the same block. The community catalogue's base policy grants /usr/local/bin/claude and /usr/bin/node access to api.anthropic.com, while a separate block allows /usr/bin/git to reach github.com for GET /**/info/refs* and POST /**/git-upload-pack (clone and fetch) with the push path commented out.[11] Host patterns accept a * wildcard inside the first DNS label or as an entire middle label, and a recursive ** only as the whole first label; they reject *, ** and TLD wildcards such as *.com, failing at load time rather than mismatching silently at the proxy.[5] A network_middlewares section can insert ordered request middleware, including a built-in regular-expression redactor, after policy admits a request but before credentials are injected.[7]

NVIDIA's September 2026 walkthrough shows the same model with access: read-only on the GitHub REST API: curl inside the sandbox can GET from api.github.com, a POST to the same host is blocked, and the replacement policy is applied with openshell policy set without restarting the sandbox.[35]

Agent-proposed policy, human approval

Because a deny-by-default policy will block legitimate work, OpenShell includes a proposal loop: when an agent hits a constraint it can reason about the denial and propose a policy update, with a human holding final approval.[2] When the policy advisor is enabled, layer-7 denial responses carry structured next_steps and an agent_guidance string so a generic agent treats the block as a step rather than a dead end.[5] In 0.1 the policy advisor is enabled per sandbox (for example with openshell settings set <sandbox-name> --key agent_policy_proposals_enabled --value true), proposals stay pending for human review by default, and the agent cannot approve its own request.[35]

Proposals are checked, not trusted. A Z3-based policy prover was merged into the repository in April 2026.[62] Every submitted chunk runs through openshell-prover, which builds a Z3 model of the merged policy plus the sandbox's attached credentials and reports categorical findings: reach granted to link-local or cloud-metadata addresses, layer-7 bypass by a credentialed non-HTTP binary such as ssh or nc, new credentialed reach for a binary, or a new HTTP method on an endpoint the binary could already reach. Auto-approval fires only if the approval mode is explicitly set to auto, the prover delta is empty, and advisory security notes are empty. The default at both scopes is manual, so with no configuration every proposal waits for a human.[5]

OpenShell 0.1.0 added a second use of the prover: a boundary check, run with the standalone openshell-prover command, that verifies a candidate policy grants no access beyond an operator-written boundary policy, covering filesystem access, process identity, Landlock settings, network connections and REST requests; for rules it cannot model, such as GraphQL or MCP rules, it reports that it cannot check the policy rather than ignoring them.[44] The New Stack called the prover "the important new component added in this update", while NVIDIA's walkthrough listed it as one of five new capabilities in 0.1.0.[63][35] Golshan told reporters it is "deterministic" and "not LLM as a judge".[63] NVIDIA says that in long-horizon adversarial experiments, frontier agents with reduced safeguards spent up to two hours trying to persuade an AI reviewer to grant permissions that would let them modify a protected GitHub repository, and that no protected repository writes occurred.[35]

The privacy router

In the 0.0.x series the privacy router decided where inference went. Sandbox code called a local HTTPS endpoint, inference.local; the router stripped the caller's Authorization header, forwarded only a per-provider allowlist of inference headers, injected credentials from a gateway-held provider record, and forwarded to the configured backend, so the agent never saw the real key. Anything sent directly to an external host such as api.openai.com was judged by the network policy instead.[8]

NVIDIA described the intent as keeping sensitive context on-device with local open models and routing to frontier models only when policy allows, driven by the operator's cost and privacy policy rather than the agent's preference.[2] Supported backends were NVIDIA, Anthropic, Google Vertex AI, AWS Bedrock through a translating bridge, and any OpenAI-compatible provider, which covered self-hosted Ollama and vLLM servers. One provider and one model defined inference for a whole gateway, so every sandbox there shared the same backend.[8][17]

OpenShell 0.1.0 removed the openshell inference commands, the route APIs and the inference.local endpoint.[40] Model access is now granted like any other service: an operator attaches a provider to the sandbox, the workload calls the provider's native API, and OpenShell evaluates the request against the network policy derived from the provider profile, substituting the real credential only at an endpoint that profile authorizes.[43]

Audit trail

Observable sandbox behaviour is logged in Open Cybersecurity Schema Framework format against vendored OCSF schemas (v1.7.0 in July 2026, v1.8.0 from August 2026): proxy decisions map to network or HTTP activity, plus SSH activity, process activity, configuration state change and detection finding classes. Forward-proxy successes are emitted only after the full chain succeeds, so an allowed record cannot coexist with a later denial for the same request, and the project forbids logging secrets, tokens or query parameters because the JSONL output is expected to be shipped elsewhere.[5][13] OpenShell 0.1.0 made the OCSF schema version configurable for compatibility with existing security information and event management (SIEM) systems.[39]

The 0.1 redesign

OpenShell 0.1.0 implemented a new sandbox architecture (RFC 0012, accepted on 10 September 2026 and implemented in a pull request merged on 16 September 2026) that moves the trusted component out of the agent's workload.[53][41] Each sandbox now has a supervisor on the trusted side of the boundary, which checks requests against policy, supplies credentials, resolves DNS and opens approved connections, and a sandbox component inside the boundary that owns the agent's processes, identifies the calling executable from trusted /proc data, and hands TCP opens and DNS queries to the supervisor using seccomp user notification.[41] The runtime grants neither the trusted components nor the agent child any Linux capability inside the workload; the agent runs as a single non-root identity with no_new_privs, Landlock and a final syscall filter.[6][41]

RuntimeWhere the supervisor runsHow direct egress is blocked
DockerIts own containerWorkload container has networking turned off
PodmanIts own containerWorkload container has networking turned off
KubernetesIts own podNetworkPolicy allows only the supervisor service
VMA process on the hostGuest has no network device

Expanded article table

Supervisor and sandbox talk over a mutually authenticated HTTP/2 connection (a Unix socket for Docker and Podman, TCP for Kubernetes, vsock for MicroVM). The agent cannot run before the supervisor confirms its controls, and if the supervisor disconnects the sandbox freezes the agent; only the same supervisor process can reconnect and resume it.[41] The 0.1 architecture also adds workspaces, an access and resource isolation boundary for multi-team use, and extension points for gateway interceptors, supervisor middleware, isolation backends and drivers for compute runtimes and secret stores.[35][65]

Installation, platforms and CLI

In the 0.0.x series OpenShell installed as a static binary, from PyPI with uv tool install -U openshell, or into Kubernetes via a Helm chart on GHCR.[48] From 0.1 the install script sets OpenShell up with Homebrew on macOS (running the gateway as a Homebrew service) and with Debian or RPM packages on Linux, with a Snap package as an option; the Helm chart is still published on GHCR.[64] Local 0.0.x installations cannot be upgraded in place: operators must remove old sandboxes, export provider profiles and uninstall before installing 0.1.0.[40]

Supported hosts are Linux on x86_64 and aarch64, macOS on Apple Silicon through Docker Desktop, and Windows through WSL 2, which the docs mark experimental. In the 0.0.x series seccomp was required and Landlock only recommended; the 0.1 support matrix requires both, Landlock at ABI 3 or newer (introduced in Linux 6.2) and seccomp with nested user-notification support. On macOS both operate inside Docker Desktop's Linux VM rather than the host kernel. Compute drivers are Docker, Podman, Kubernetes and MicroVM backed by KVM or Hypervisor.framework.[9][42] In June 2026 NVIDIA said it was collaborating with Microsoft to bring OpenShell to Windows built on Microsoft eXecution Containers (MXC); as of September 2026 the OpenShell documentation lists a Windows MXC runtime as "coming soon".[50][51] Under the 0.1 release policy, stable releases generally ship weekly, and security and critical reliability fixes are provided for the latest and previous minor release lines.[42]

The CLI is organised around gateways, sandboxes, providers and policies. In the 0.0.x README openshell sandbox create -- claude was the canonical first command, and openshell term opens a k9s-style terminal dashboard.[1][12] The 0.1 README starts from openshell sandbox create --name demo with a minimal Ubuntu image that has no agent installed, and its first-agent tutorial runs OpenCode against an OpenRouter model.[47] GPU passthrough was available with --gpu and marked experimental in the 0.0.x series; it required host NVIDIA drivers and the NVIDIA Container Toolkit, and the default base image shipped without GPU libraries.[48] NVIDIA's launch post advertised openshell sandbox create --remote spark --from openclaw for running on an NVIDIA DGX Spark; the July 2026 CLI reference had no --remote flag on sandbox create, and remote gateways were registered with openshell gateway add --remote <USER@HOST> instead.[2][12] NVIDIA lists DGX Spark, DGX Station and RTX-equipped PCs as target hardware.[2]

The 0.1 README also offers public agent skills, installed with npx skills add NVIDIA/OpenShell, that teach a coding agent to drive the OpenShell CLI, write sandbox policies and debug gateways and inference routing, plus SDKs for Python, TypeScript, Go and Rust that connect applications to a gateway.[47] NVIDIA's 28 September release said OpenShell and skills are available through the NVIDIA developer resources page and GitHub.[36] OpenShell collects anonymous operational telemetry by default; it can be disabled with OPENSHELL_TELEMETRY_ENABLED=false on the gateway or compiled out entirely.[47]

NemoClaw, OpenClaw and Hermes Agent

OpenShell is the runtime layer; NemoClaw is the blueprint that assembles a model, an agent harness and that runtime into something installable. In July 2026 NVIDIA described NemoClaw as "an open source reference stack for running always-on AI agents more safely inside NVIDIA OpenShell sandboxes"; by September its README said "supported AI agents" instead, and listed guided onboarding, managed inference, network policy, managed integrations, snapshots and lifecycle operations through the NemoClaw CLI. NemoClaw is also Apache 2.0 and still calls itself an alpha project. Its default agent is OpenClaw, the always-on assistant framework; alternatives are Hermes Agent and LangChain Deep Agents Code.[18]

Hermes Agent is Nous Research's open harness. In OpenShell's 0.0.x supported-agent table it appeared alongside OpenClaw as an agent sourced through NemoClaw, meaning routing and policy came from the NemoClaw blueprint rather than the base sandbox image; Claude Code, opencode, OpenAI Codex and GitHub Copilot CLI were preinstalled in the community base image, alongside community sandboxes for Ollama and Pi.[48] OpenShell 0.1.0 replaced that fallback with a minimal Ubuntu image containing no agent CLI and no image-baked policy, so operators build images with the agents they need.[40] NVIDIA's 0.1.0 walkthrough says the runtime supports Codex, Claude Code, Pi, Hermes "and future frameworks".[35] Cursor was named among compatible agents in the March launch post but was absent from the repository's supported-agent table, appearing instead as a remote-editor option (--editor cursor).[2][12]

Comparison with other agent sandboxing approaches

ApproachIsolation mechanismWhere it runsNotes
NVIDIA OpenShellLandlock, seccomp, per-binary L7 proxy; from 0.1 a supervisor outside the workload, with container, pod or MicroVM boundariesSelf-hostedDeny-by-default egress bound to the calling binary; credential substitution outside the workload; formal policy prover; OCSF audit log; Apache 2.0[1][5][6][41]
Docker / OCI containersNamespaces and cgroups over a shared host kernelAnywhereThe substrate OpenShell's Docker driver builds on, not an agent policy layer itself[10]
gVisorUserspace application kernel (Sentry) intercepting syscalls, plus a Gofer for filesystem accessDocker, Kubernetes, or directly through the runsc OCI runtimePositions itself as a third approach between syscall filters and virtual machines, with many of a VM's security benefits and a smaller footprint[28]
FirecrackerKVM microVMsSelf-managed hosts; AWS Lambda MicroVMs as a managed serviceApache 2.0; under 125 ms startup and under 5 MiB memory footprint per microVM[29]
E2BCloud sandbox service for AI-generated codeHostedApache 2.0; continuous runs of up to 1 hour on the Base tier and 24 hours on Pro, with pause and resume[31][56]
Modal SandboxesManaged containers for untrusted user or agent codeHostedFive-minute default lifetime, adjustable up to 24 hours; readiness probes, custom images[30]
DaytonaHosted sandbox infrastructure for AI-generated codeHostedClaims sub-90 ms sandbox creation; its public GitHub repository stopped receiving updates when core development moved to a private codebase in June 2026[32][57]
Anthropic sandbox-runtimesandbox-exec on macOS, bubblewrap with network namespaces on Linux, a dedicated user account with a Windows Filtering Platform fence on Windows; proxy-based domain allowlistLocal, no containerBeta research preview developed for Claude Code; can wrap agents, local MCP servers, bash commands and arbitrary processes[26]
OpenAI Codex CLIsandbox_mode of read-only, workspace-write or danger-full-access, with protected .git and .codex pathsLocalApproval and sandbox settings combined per session[27]

Expanded article table

The layers do different jobs. Docker, gVisor and Firecracker are isolation substrates with different strength and cost trade-offs; OpenShell builds on containers, Kubernetes pods or microVMs through its driver interface rather than replacing them. E2B, Modal and Daytona are hosted execution services consumed as an API by an agent framework. Anthropic's sandbox-runtime and Codex's sandbox modes are local controls built alongside a specific coding agent. OpenShell's claim is to be the horizontal layer beneath all of them: agent-agnostic, self-hosted, enforcing one policy the agent cannot reach.[15][1] An August 2026 NVIDIA post on the agent stack placed OpenShell in a "secure runtime" layer responsible for isolation, identity, policy, credentials and audit, beneath agent harnesses such as Claude Code, Codex, Hermes and Pi and above inference infrastructure such as NVIDIA Dynamo.[49]

Open Agent Safety Platform

On 28 September 2026 NVIDIA announced the NVIDIA Open Agent Safety Platform, "an open software platform and reference system design" made up of OpenShell and the NVIDIA Sentry reference system design.[36] The press release said: "Now broadly available, OpenShell provides a secure runtime boundary for controlling how autonomous AI agents execute tasks across open and closed models."[36] It came after OpenAI, Anthropic, Meta and Google disclosed incidents in which their models escaped evaluation sandboxes; CNBC reported that an NVIDIA representative told reporters the platform could have prevented the July OpenAI-Hugging Face agent incident.[38][63] Jensen Huang described the platform on CNBC's Squawk Box as "a browser for agents".[38]

NVIDIA says OpenShell delivers its protection "with minimal overhead" on NVIDIA Vera, which the company calls "the first purpose-built CPU for agentic AI", and that as open source software OpenShell "can also be extended to work with third-party compute platforms, including those from Arm and Intel".[36] Sentry is an out-of-band watchdog on BlueField-4 DPUs, built on NVIDIA DOCA software; NVIDIA says it quarantines an agent that moves outside its software boundary "in milliseconds".[36] NVIDIA's technical blog describes Sentry as "an optional security layer alongside OpenShell" for organizations that want an additional, independent layer, with DOCA connecting the BlueField security foundation to OpenShell policy, and says that on a Vera system with BlueField-4 enabling these protections "is just a software update".[37] Justin Boitano, NVIDIA's vice president of enterprise AI, said in a press briefing reported by The New Stack that unlike OpenShell, Sentry is not open source, though it has open APIs, and that "the DPU is really optional in these architectures".[63]

NVIDIA said more than 100 organizations were working with the platform's technologies.[36] The integrations NVIDIA described individually, most of which name OpenShell, are:

OrganizationOpenShell integration, as described by NVIDIA on 28 September 2026
AnthropicClaude Managed Agents run the agent loop on a separate server from the sandboxes where work executes; integrations with OpenShell and BlueField let enterprises enforce control over agent access through those sandboxes[36]
SalesforceIntegrated OpenShell with Slack, so teams can view agent activity and audit events and approve or reject agent requests for additional permissions from Slack[36]
SAPEmbedding OpenShell with the Joule Studio runtime, part of the SAP Business AI Platform; also contributing engineering work to OpenShell[36]
Red HatRuns OpenShell and DOCA on Red Hat AI Factory with NVIDIA; Canonical, SUSE and Red Hat are integrating the platform into their operating systems[36]
SpaceXAIUsing the platform for Cursor coding agents and Grok models[36]
Scale AIIncorporating platform technologies into the agentic infrastructure layer of Scale GenAI Portfolio[36]
CadenceUses OpenShell with its ChipStack Autonomous RTL Design Engineer for chip design[35]
Gecko RoboticsUses OpenShell to govern agents making decisions on physical robots; NVIDIA also names Figure and Skild AI among robotics companies building with OpenShell[35][36]

Expanded article table

Cursor was acquired by SpaceX on 14 August 2026, completing a process that Cursor said began with its April 2026 partnership with SpaceXAI.[58] Mike Nicolls, president at SpaceXAI, said in the release that "safety should be enforced outside the model by additional controls the agent can't get past."[36] NVIDIA's walkthrough separately says Slack is building an on-demand agent platform on OpenShell.[35] SAP is also working with NVIDIA on interoperability standards through the Open Secure AI Alliance, a separate NVIDIA-initiated group of more than 120 organizations that NVIDIA's release describes as governed by the Linux Foundation.[36]

Adoption and reception

As of 28 September 2026 the repository had 8,960 stars, 1,298 forks and 124 listed contributors, and GitHub counted about 1,260 issues and 2,200 pull requests.[47] Coverage at launch came from MarkTechPost, which highlighted the per-binary, per-endpoint, per-method policy model and the fact that OpenShell is agent-agnostic rather than requiring a rewrite into a particular SDK.[16] Community discussion runs on GitHub and in an #openshell-dev channel on the CNCF Slack workspace.[35][47]

LangChain announced on the day of the GTC keynote an enterprise agent platform built with NVIDIA that incorporates OpenShell, and said the collaboration lays the groundwork for Deep Agents to operate within GPU-accelerated compute sandboxes; The New Stack later reported that LangChain would contribute openly to the OpenShell repository.[20][15]

An early enterprise adopter was ServiceNow, with Project Arc, announced on 5 May 2026 at ServiceNow Knowledge 2026 in Las Vegas with Jensen Huang and ServiceNow chief executive Bill McDermott on stage together. Project Arc is a long-running autonomous desktop agent for knowledge workers that uses OpenShell as its secure runtime, connecting to ServiceNow Action Fabric for workflow context and AI Control Tower for oversight, with ServiceNow contributing back to the project. Jon Sigler, ServiceNow's executive vice president and general manager for AI Platform, said the combination delivers "the governance and security that enterprise AI requires".[19][15] The partnership also advances NOWAI-Bench, an open benchmarking suite for enterprise agents built on NVIDIA's NeMo Gym library, whose EnterpriseOps-Gym component NVIDIA says Nemotron 3 Super currently leads among open models.[19] In March 2026 NVIDIA named Cisco, CrowdStrike, Google Cloud, Microsoft Security and TrendAI as security partners, said OpenShell would run on Canonical Ubuntu, Microsoft Windows and Red Hat OpenShift, and said it would be integrated into SAP and ServiceNow platforms.[14]

Limitations and criticism

Until the 0.1.0 release in September 2026 the project labelled itself alpha and "single-player mode": one developer, one environment, one gateway, with multi-tenant deployment described as a goal.[48] Several specific gaps were documented in the open tracker, and several of those issues were closed by September 2026.

A researcher demonstrated in May 2026 that DNS was a covert exfiltration channel: with no network policy configured, an HTTP request to an attacker-controlled domain was correctly denied by the proxy, but the DNS lookup for that hostname had already carried data encoded in subdomain labels to the attacker's nameserver. Russell Bryant, whose GitHub profile lists Red Hat, posted a fix that removed a DNS lookup for a hostname denied by policy (pull request #1329), and the issue was closed on 15 May 2026.[21][54][61] A proxy that mediates TCP does not automatically mediate name resolution; in the 0.1 design DNS queries are also handed to the supervisor.[41]

Kubernetes deployment was constrained by the privilege the in-pod supervisor needed. Sandbox pods requested CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_SYS_PTRACE, CAP_SYSLOG and runAsUser: 0, which Red Hat OpenShift's default restricted-v2 SecurityContextConstraint forbids; the filed issue argued that granting a custom constraint with those capabilities weakens the cluster and is "a non-starter for many enterprise deployments".[22] A related draft RFC from external contributors proposed splitting supervisor and agent into separate pods, the agent running with zero capabilities as non-root, optionally under a gVisor or Kata RuntimeClass, because the shared-pod design "creates a wide blast radius if the agent escapes its confinement".[23] In the same RFC thread, one commenter argued that identifying "which binary initiated the connection" could already be defeated with LD_PRELOAD.[23] The SCC issue was closed as completed on 15 August 2026 and the split-pod RFC was closed as not planned on 21 September 2026, after the RFC 0012 architecture had been merged.[22][23][53] The 0.1 OpenShift guide states that OpenShell does not require the privileged SCC or any added Linux capability.[45]

High availability was also limited. In July 2026 reliable readiness required a single gateway replica because supervisor sessions were process-local, with the fix deferred to #1868.[10] That change, "support HA gateway rebalancing", was merged on 21 September 2026, and the 0.1 documentation describes running two or more gateway replicas that share state through PostgreSQL.[52][46] The July 2026 default-policy reference covered Claude Code fully, OpenCode partially and Codex not at all, so non-Claude agents needed a custom policy naming their endpoints and binaries.[59] In 0.1 the fallback policy denies all network access and attached providers contribute their own network rules.[40][43]

Other rough edges remain documented: Windows runs only through WSL 2 until the MXC runtime ships; 0.1.0 introduced breaking changes across deployments, policies, APIs and SDKs, and mixed 0.0.x and 0.1.0 components are not supported.[42][51][40] The prover's guarantees apply only to the policy features its model represents, which NVIDIA says it is extending.[44]

Significance

OpenShell is an argument as much as a product: that the security primitives for AI agents belong below the application layer, where the model cannot reason its way past them, and that this layer should be shared open infrastructure rather than a per-vendor feature. Comparable enforcement exists in narrower form inside individual coding agents, while OpenShell tries to make it agent-agnostic, self-hostable and auditable in one package. Whether it becomes the common layer depends on whether harness vendors, cloud platforms and governance tools converge on it. The ServiceNow and LangChain commitments in the first half of 2026, and the Anthropic, Salesforce, SAP and SpaceXAI integrations announced with the Open Agent Safety Platform in September 2026, are NVIDIA's evidence that they are starting to.[19][20][36]

See also

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7NVIDIA, "OpenShell: the safe, private runtime for autonomous AI agents" (README and repository metadata), GitHub, accessed 31 July 2026. github.com/...OpenShell
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12Ali Golshan, Alex Watson and John Myers, "Run Autonomous, Self-Evolving Agents More Safely with NVIDIA OpenShell", NVIDIA Technical Blog, 16 March 2026. developer.nvidia.com/...fely-with-nvidia-openshell
  3. ^NVIDIA, "NVIDIA CEO Jensen Huang and Global Technology Leaders to Showcase Age of AI at GTC 2026", NVIDIA Newsroom, 3 March 2026. nvidianews.nvidia.com/...ase-age-of-ai-at-gtc-2026
  4. ^NVIDIA, "Overview of NVIDIA OpenShell", NVIDIA OpenShell documentation, accessed 31 July 2026. docs.nvidia.com/...overview
  5. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9NVIDIA, "Security Policy" (architecture/security-policy.md), NVIDIA/OpenShell, GitHub, accessed 31 July 2026. github.com/...security-policy.md
  6. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8NVIDIA, "Sandbox" (architecture/sandbox.md), NVIDIA/OpenShell, GitHub, accessed 31 July 2026 and 28 September 2026. github.com/...sandbox.md
  7. ^1 ^2 ^3 ^4 ^5NVIDIA, "Customize Sandbox Policies", NVIDIA OpenShell documentation, accessed 31 July 2026 (page retired in the 0.1 documentation restructure). docs.nvidia.com/...policies
  8. ^1 ^2NVIDIA, "Inference Routing", NVIDIA OpenShell documentation, accessed 31 July 2026 (page retired in the 0.1 documentation restructure). docs.nvidia.com/...inference-routing
  9. ^NVIDIA, "Support Matrix", NVIDIA OpenShell documentation, accessed 31 July 2026 (moved in the 0.1 documentation restructure). docs.nvidia.com/...support-matrix
  10. ^1 ^2NVIDIA, "Compute Runtimes" (architecture/compute-runtimes.md at commit c42268b, 31 July 2026), NVIDIA/OpenShell, GitHub. github.com/...compute-runtimes.md
  11. ^1 ^2NVIDIA, "sandboxes/base/policy.yaml", NVIDIA/OpenShell-Community, GitHub, accessed 31 July 2026. github.com/...policy.yaml
  12. ^1 ^2 ^3NVIDIA, "OpenShell CLI Reference" (.agents/skills/openshell-cli/cli-reference.md at commit c42268b, 31 July 2026; removed from main in September 2026), NVIDIA/OpenShell, GitHub. github.com/...cli-reference.md
  13. ^NVIDIA, "Vendored OCSF Schemas" (crates/openshell-ocsf/schemas/ocsf/README.md at commit c42268b, 31 July 2026; updated to OCSF v1.8.0 on 18 August 2026), NVIDIA/OpenShell, GitHub. github.com/...README.md
  14. ^1 ^2Ali Golshan, "How Autonomous AI Agents Become Secure by Design With NVIDIA OpenShell", NVIDIA Blog, 23 March 2026. blogs.nvidia.com/...autonomous-ai-agents-openshell
  15. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Darryl K. Taft, "Jensen Huang and Bill McDermott bet on OpenShell to secure enterprise AI agents", The New Stack, 12 May 2026. thenewstack.io/nvidia-openshell-agent-runtime
  16. ^Asif Razzaq, "NVIDIA AI Open-Sources 'OpenShell': A Secure Runtime Environment for Autonomous AI Agents", MarkTechPost, 18 March 2026. marktechpost.com/...nment-for-autonomous-ai-agents
  17. ^NVIDIA, "OpenShell", NVIDIA Perspectives, accessed 28 September 2026 (last updated 17 September 2026). perspectives.nvidia.com/nvidia-openshell
  18. ^NVIDIA, "NemoClaw" (README), GitHub, accessed 31 July 2026 and 28 September 2026. github.com/...NemoClaw
  19. ^1 ^2 ^3NVIDIA, "NVIDIA and ServiceNow Partner on New Autonomous AI Agents for Enterprises", NVIDIA Blog, 5 May 2026. blogs.nvidia.com/...tonomous-ai-agents-enterprises
  20. ^1 ^2 ^3LangChain, "LangChain Announces Enterprise Agentic AI Platform Built with NVIDIA", LangChain blog, 16 March 2026. langchain.com/...nvidia-enterprise
  21. ^"DNS-based data exfiltration bypasses sandbox network policy", issue #1169, NVIDIA/OpenShell, GitHub, opened 5 May 2026, closed 15 May 2026. github.com/...1169
  22. ^1 ^2"feat: Support restricted SecurityContextConstraints for managed Kubernetes platforms", issue #899, NVIDIA/OpenShell, GitHub, opened 20 April 2026, closed 15 August 2026. github.com/...899
  23. ^1 ^2 ^3"Proposal: Split Supervisor and Agent into Separate Pods with gVisor Isolation", issue #981, NVIDIA/OpenShell, GitHub, opened 26 April 2026, closed 21 September 2026. github.com/...981
  24. ^Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz and Mario Fritz, "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection", arXiv:2302.12173, 23 February 2023. arxiv.org/...2302.12173
  25. ^OWASP GenAI Security Project, "OWASP Top 10 for LLM Applications 2025", accessed 31 July 2026. genai.owasp.org/llm-top-10
  26. ^Anthropic, "sandbox-runtime" (README), GitHub, accessed 28 September 2026. github.com/...sandbox-runtime
  27. ^OpenAI, "Codex configuration: sandbox and approvals", Codex documentation, accessed 31 July 2026. learn.chatgpt.com/...config-basic
  28. ^The gVisor Authors, "What is gVisor?", gVisor documentation, accessed 28 September 2026. gvisor.dev/docs
  29. ^Firecracker project, "Secure and fast microVMs for serverless computing", accessed 28 September 2026. firecracker-microvm.github.io
  30. ^Modal, "Sandboxes", Modal documentation, accessed 28 September 2026. modal.com/...sandbox
  31. ^E2B, "E2B: secure open source cloud runtime for AI apps and agents" (repository), GitHub, accessed 31 July 2026. github.com/...E2B
  32. ^Daytona, "Secure and Elastic Infrastructure for Running Your AI-Generated Code", accessed 28 September 2026. daytona.io
  33. ^1 ^2NVIDIA, "Ali Golshan" (author profile), NVIDIA Technical Blog, accessed 28 September 2026. developer.nvidia.com/...aligolshan
  34. ^Kyle Wiggers, "Nvidia reportedly acquires synthetic data startup Gretel", TechCrunch, 19 March 2025. techcrunch.com/...es-synthetic-data-startup-gretel
  35. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12Alex Watson and Ali Golshan, "Add Runtime Controls to AI Agents with NVIDIA OpenShell", NVIDIA Technical Blog, 28 September 2026. developer.nvidia.com/...ents-with-nvidia-openshell
  36. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17NVIDIA, "NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment", NVIDIA Newsroom, 28 September 2026. nvidianews.nvidia.com/...open-agent-safety-platform
  37. ^1 ^2 ^3John Myers, Alex Watson, Ali Golshan and Ofir Arkin, "NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring", NVIDIA Technical Blog, 28 September 2026. developer.nvidia.com/...n-silicon-agent-monitoring
  38. ^1 ^2Kif Leswing, "Nvidia releases software platform to stop AI agents from misbehaving", CNBC, 28 September 2026. cnbc.com/...nvidia-releases
  39. ^1 ^2NVIDIA, "OpenShell v0.1.0" (release notes), NVIDIA/OpenShell, GitHub, 25 September 2026. github.com/...v0.1.0
  40. ^1 ^2 ^3 ^4 ^5 ^6NVIDIA, "Upgrade to NVIDIA OpenShell 0.1.0", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...0-1-0
  41. ^1 ^2 ^3 ^4 ^5 ^6NVIDIA, "Architecture", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...architecture
  42. ^1 ^2 ^3NVIDIA, "Support Matrix", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...support-matrix
  43. ^1 ^2NVIDIA, "Inference", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...inference
  44. ^1 ^2NVIDIA, "Policy Prover", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...prover
  45. ^NVIDIA, "OpenShift", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...openshift
  46. ^NVIDIA, "High Availability", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...high-availability
  47. ^1 ^2 ^3 ^4 ^5NVIDIA, OpenShell README and repository metadata, NVIDIA/OpenShell, GitHub, accessed 28 September 2026. github.com/...README.md
  48. ^1 ^2 ^3 ^4NVIDIA, OpenShell README at commit c42268b (31 July 2026), NVIDIA/OpenShell, GitHub. github.com/...README.md
  49. ^Johnny Greco, Kirit Thadaka, Ali Golshan and Alex Watson, "Where Security Fits in an AI Agent Stack", NVIDIA Technical Blog, 21 August 2026. developer.nvidia.com/...-fits-in-an-ai-agent-stack
  50. ^Annamalai Chockalingam and Gerardo Delgado, "Build Personal AI Agents on Windows PCs with New Tools from Microsoft and NVIDIA", NVIDIA Technical Blog, 2 June 2026. developer.nvidia.com/...-from-microsoft-and-nvidia
  51. ^1 ^2NVIDIA, "Sandbox Runtimes", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...runtimes
  52. ^"feat(kubernetes): support HA gateway rebalancing", pull request #1868, NVIDIA/OpenShell, GitHub, merged 21 September 2026. github.com/...1868
  53. ^1 ^2"feat(isolation): implement the RFC 0012 sandbox architecture", pull request #2942, NVIDIA/OpenShell, GitHub, merged 16 September 2026. github.com/...2942
  54. ^Russell Bryant (russellb), GitHub profile, accessed 28 September 2026. github.com/russellb
  55. ^OWASP GenAI Security Project, "LLM01:2025 Prompt Injection", accessed 28 September 2026. genai.owasp.org/...llm01-prompt-injection
  56. ^E2B, "Sandbox lifecycle", E2B Docs, accessed 28 September 2026. e2b.dev/...sandbox
  57. ^Daytona, "daytonaio/daytona" (README), GitHub, accessed 28 September 2026. github.com/...daytona
  58. ^Cursor Team, "Cursor is now a part of SpaceX", Cursor blog, 14 August 2026. cursor.com/...joining-spacex
  59. ^NVIDIA, "Default Policy Reference" (docs/reference/default-policy.mdx at commit c42268b, 31 July 2026), NVIDIA/OpenShell, GitHub. github.com/...default-policy.mdx
  60. ^NVIDIA, "OpenShell v0.1.2" (release notes), NVIDIA/OpenShell, GitHub, 28 September 2026. github.com/...v0.1.2
  61. ^"fix(sandbox): remove DNS resolution from mechanistic mapper to prevent data exfiltration", pull request #1329, NVIDIA/OpenShell, GitHub, merged 15 May 2026. github.com/...1329
  62. ^"feat(prover): add native Rust policy prover with Z3 solver", pull request #741, NVIDIA/OpenShell, GitHub, merged 9 April 2026. github.com/...741
  63. ^1 ^2 ^3 ^4Frederic Lardinois, "Nvidia launches Open Agent Safety Platform to lock down rogue AI agents", The New Stack, 28 September 2026. thenewstack.io/nvidia-openshell-sentry-agents
  64. ^NVIDIA, "Installation", NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...installation
  65. ^NVIDIA, "Extensibility" (overview), NVIDIA OpenShell documentation, accessed 28 September 2026. docs.nvidia.com/...overview

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

3 revisions · v4 · 6,513 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Full independent verification 28 Sep 2026 (xg11 V3); v4 adds only an internal link to nvidia_agent_toolkit

Cite this page: AI Wiki. "NVIDIA OpenShell." aiwiki.ai, updated 28 Sept 2026, fact-checked 28 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/nvidia_openshell

Suggest edit