OpenAI Agents API Public Beta: Managed Sandboxes vs Self-Hosted Agent Runtimes
The most expensive line item in a production agent is rarely the model. It is the box the agent runs code in, and that box has a habit of staying alive long after the useful work has finished. The OpenAI Agents API entered public beta in September 2026 with a pitch that sounds almost too tidy: describe the task, pick a model, name the tools, and OpenAI runs the whole loop, including the sandbox, on its own infrastructure. The widely quoted headline number is about $0.03 for a 20-minute sandbox session.
That number is real, but it is the smallest tier, and it is only one term in a cost equation that also includes tokens, idle time, minimum billing, and the engineering you do not have to build. This post takes the managed-versus-self-hosted decision apart: what OpenAI has actually published, what it has not, how the isolation and cost models compare against Firecracker and gVisor runtimes you operate yourself, and a decision framework you can apply to your own workload.
What this covers: the confirmed facts about the launch, the architecture of a managed agent runtime, the sandbox pricing tiers and a clearly labelled illustrative cost model, a decision matrix against self-hosted alternatives, security and failure modes, and a short checklist for choosing and piloting.
Context and Background
An AI agent, in the operational sense used here, is a loop: a model reads state, decides on an action, a tool executes that action, and the result feeds the next model call. The tools that matter most for general-purpose agents are the ones that run arbitrary code, because code execution turns a language model from a text generator into something that can read files, transform data, run tests, and produce artifacts. Arbitrary code from a model is untrusted code by definition, which is why every serious agent platform needs a sandbox.
Until 2026 most teams assembled this themselves. They picked an orchestration library, wired a model client, and bolted on an execution environment, either a container on their own cluster, a micro-virtual-machine (microVM) service, or a third-party sandbox vendor. The orchestration side is covered in our benchmark of LangGraph, the OpenAI agent tooling and Google ADK, and the isolation side in our deep dive on Firecracker, gVisor and Kata for agent sandboxes. The Agents API collapses much of that assembly into one managed surface.
The competitive context matters. Anthropic already sells a managed agent runtime, and independent cost write-ups published in the weeks after the OpenAI launch compare the two on a per-session-hour basis. Sandbox specialists such as E2B, Daytona, Modal and Vercel have been selling isolated execution as a product for some time, and OpenAI’s announcement lists several of them as integration partners rather than competitors. That detail tells you how OpenAI sees the market: the model and the harness are the product, and the sandbox is a pluggable resource.
For primary sources, start with OpenAI’s own Introducing the Agents API post and the developer community announcement thread. Everything below that is a vendor claim is attributed to one of them, and everything that comes from third-party analysis is labelled as such.
What OpenAI Actually Shipped
The short answer: the OpenAI Agents API is a managed service, announced on 10 September 2026 and available to all developers in public beta, that runs agents built on OpenAI’s Codex harness. You specify a task, model, tools and environment in one call. OpenAI hosts the sandbox, or you connect your own, and you pay for tokens and tools rather than a platform fee.
A note on dates, because the launch is often misdated in roundups. OpenAI’s announcement and the developer community thread are dated 10 September 2026, not the middle of the month. Treat later dates in secondary coverage as the date the commentary appeared, not the date the product shipped.
Here is what the primary sources state. The service lets you “build and run cloud agents with the Codex harness, fully managed by OpenAI.” Tool support covers the Model Context Protocol (MCP), custom functions and built-in web search. The harness provides automatic context compaction for long sessions, tool search to reduce token usage, programmatic parallel tool calling, and subagent orchestration, with a configurable max_concurrent_subagents setting that the announcement shows at 3. OpenAI states there is no separate fee for the Agents API itself: “you simply pay for the tokens and tools your agents use.”
On environments, the community announcement says agents can run code, work with files and produce artifacts in OpenAI-hosted sandboxes, billed at “standard container rates” with model usage billed separately. It also describes the option to “bring your own sandbox” or connect a provider, with CPU, GPU and memory options, and either fully managed environments or deployments inside your own virtual private cloud (VPC). Named integration partners are Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop and Vercel.
What is not disclosed
The gaps are as informative as the features. In the sources fetched for this article, OpenAI does not detail the isolation technology behind its hosted sandboxes, beyond stating the infrastructure mirrors what Codex and ChatGPT use. It does not state maximum session length, data retention for sandbox contents, regional availability, or how many sandboxes a subagent fan-out may allocate. A third-party pricing analysis lists the same unknowns: which container tier applies by default, whether idle time is billed, maximum session length, and concurrent subagent sandbox allocation.
For a prototype this is acceptable. For a regulated workload it is a procurement blocker until answered, and the right move is to ask OpenAI in writing rather than infer from the Codex product.
Beta means beta
“Public beta” with no stated general-availability date carries the usual meaning: interfaces can change, quotas can move, and service-level commitments are usually absent. The practical consequence is architectural. If you build on the managed runtime, put a thin adapter between your application and the API so that a changed parameter or a repriced tier costs you an afternoon, not a quarter.
Anatomy of a managed agent runtime
It helps to separate four planes that every agent runtime has, whether you buy it or build it. The control plane accepts a task and holds the configuration. The reasoning plane is the model-calling loop, including context management. The tool plane brokers calls to MCP servers, functions and search. The execution plane is the sandbox where generated code actually runs. A managed service takes all four off your hands; a self-hosted design lets you keep any subset.

Figure 1: The four planes of an agent runtime. The Agents API manages all of them, and the sandbox plane can be swapped for a partner provider or your own VPC deployment.
The figure shows why the sandbox is the natural seam. The control, reasoning and tool planes are tightly coupled to the harness and the model, so splitting them across vendors creates latency and debugging pain. The execution plane, by contrast, speaks a narrow interface: run this command, read or write this file, return the output. That narrow interface is what lets OpenAI support nine partner providers and a bring-your-own option without redesigning the harness.
The harness is the differentiator, not the container
It is tempting to read the launch as OpenAI selling cheap containers. The container is the least differentiated piece. What a team inherits is the Codex harness behaviour: compaction that summarises older turns so long sessions do not exhaust the context window, tool search that loads tool definitions on demand instead of stuffing every schema into the prompt, and parallel tool calls issued programmatically rather than one round trip at a time.
Those features attack the two real costs of long-running agents, which are tokens and wall-clock time. Context compaction in particular changes the economics: without it, a four-hour coding session either fails when the window fills or burns tokens re-reading its own history. We discuss the underlying technique in our guide to context engineering for production LLM agents.
The honest caveat is that you cannot inspect or tune this harness the way you can an open-source orchestration library. If compaction drops a detail your workflow needs, you get the knobs OpenAI exposes and no more. That is the control you trade for convenience, and it is the first axis of the decision.
Subagents and the sandbox multiplier
Subagent support is where sandbox economics get interesting. The announcement shows a default-style limit of three concurrent subagents. If each subagent receives its own sandbox, a three-way fan-out triples container spend for the duration of the parallel phase. OpenAI has not published whether subagents share the parent’s container or get separate ones, so any cost model has to carry both cases.
This matters because parallelism is sold as a latency win. Splitting a task three ways may cut wall-clock time by half, but if every branch pays for a 4 GB container with a five-minute minimum, the bill rises faster than the speedup. Measure before you adopt fan-out as a default.
Sandbox Pricing, Billing Rules and an Illustrative Cost Model
OpenAI’s pricing page, as summarised by a third-party tracker that captured it between 9 and 12 September 2026, lists container rates per 20-minute session by memory tier. These figures come from that secondary summary of the pricing page, and OpenAI can change them during the beta, so confirm them before you budget.
| Container memory | Per 20-minute session | Equivalent per hour |
|---|---|---|
| 1 GB | $0.03 | $0.09 |
| 4 GB | $0.12 | $0.36 |
| 16 GB | $0.48 | $1.44 |
| 64 GB | $1.92 | $5.76 |
The pricing footnote, as quoted by the same tracker, says eligible container sessions are “billed by the minute, with a 5-minute minimum per session.” Two consequences follow. The rate scales linearly with memory, so the $0.03 headline applies only to the 1 GB tier. And a five-minute minimum means a workload made of many tiny tool calls, each in a fresh session, pays for five minutes every time.

Figure 2: How sandbox cost responds to memory tier, billable duration and idle time. Memory and duration multiply; the five-minute floor dominates short sessions.
Tokens still dominate at small scale, sandboxes at large
The tracker’s worked example, which is the author’s own calculation from published model rates, compares a one-hour session with 50,000 input and 15,000 output tokens. On a small, cheap model with a 1 GB sandbox, tokens come to about $0.028 and the sandbox to $0.09, so the sandbox is roughly three quarters of the total. On a 4 GB tier the sandbox share rises above 90 percent. With a frontier model on a competing managed runtime, the same hour is mostly tokens.
The lesson is not that one vendor is cheaper. It is that the dominant cost term depends on which model you pair with the container. A cheap model on a large sandbox is a sandbox-cost problem. An expensive model on a small sandbox is a token-cost problem. Optimise the one that is actually large for you.
An illustrative cost model for a hosted fleet
The following arithmetic is illustrative, not a benchmark. Assume a workload of 10,000 agent sessions per month, each running 12 minutes of real sandbox time on the 4 GB tier. Under per-minute billing with a five-minute minimum, 12 minutes is billed as 12 minutes, so each session costs 12 times the per-minute rate. At $0.12 per 20 minutes the per-minute rate is $0.006, so a session costs $0.072 and the fleet costs about $720 per month in sandbox fees.
Now change one assumption: the agent finishes its useful work at minute 4 but the session is held open for the human to review the output, and the platform bills that hold time as running. If the average hold is 20 minutes, billable time becomes 32 minutes per session and the fleet cost jumps to about $1,920. Same work, 2.7 times the bill. Whether OpenAI bills idle time is one of the undisclosed items, which is why it belongs at the top of your pre-pilot questions.
For comparison, a third-party breakdown lists E2B pricing of roughly $0.0504 per vCPU-hour plus $0.0162 per GiB-hour, which for a 2 vCPU, 4 GiB sandbox works out to about $0.166 per hour, or roughly $0.055 per 20 minutes. Daytona and Blaxel are reported at a similar level. These are provider list prices as reported by that analysis, and they differ in what a vCPU and a session include, so treat the 2x gap with OpenAI’s $0.12 per 20 minutes as a prompt to benchmark, not a verdict.
Self-Hosted Runtimes: What You Take On
Self-hosting an agent runtime means you choose an orchestrator, run the model-calling loop yourself against whichever model API you like, and operate the execution plane. The execution plane is where the real work lives, so it deserves precision about the isolation options.
A standard container shares the host kernel. A kernel vulnerability or a misconfigured capability can let code in the container reach the host, which is acceptable for code you wrote and uncomfortable for code a model generated after reading an attacker-controlled web page. Stronger options interpose a boundary. Firecracker, the microVM monitor AWS built for Lambda and Fargate, runs each workload in its own lightweight virtual machine with a minimal device model, so an escape has to defeat the hypervisor rather than a shared kernel. gVisor takes a different route: a user-space kernel, the Sentry, intercepts system calls so the application never talks to the host kernel directly. Kata Containers wraps containers in lightweight VMs behind the standard container interface. Our comparison of Firecracker, gVisor and Kata covers startup, density and syscall-compatibility trade-offs in detail.
The hidden work list
Choosing Firecracker or gVisor is the start. A production execution plane also needs a scheduler that places sandboxes on hosts, a warm pool so sessions start in a second rather than thirty, snapshotting or filesystem layering for fast restore, an egress control layer that restricts which hosts the sandbox can reach, a secrets broker that injects credentials without exposing them to generated code, log and artifact capture, and a reaper that kills runaway or orphaned sandboxes.
None of these is exotic. Together they are a platform team’s quarter of work, followed by on-call. Sandbox vendors exist because that bundle is not a differentiator for most companies, and OpenAI’s partner list is an acknowledgement of the same fact.

Figure 3: A self-hosted agent runtime. The orchestrator talks to a sandbox scheduler, which draws from a warm microVM pool behind an egress proxy and a secrets broker.
The diagram is deliberately small, and still it has seven moving parts you own. Each one is a place a failure can hide, from a warm pool that drains under a burst to a secrets broker whose misconfiguration leaks a token into a model’s context.
An illustrative self-hosted cost model
Again, this is illustrative arithmetic, not a measurement. Suppose you run Firecracker on a bare-metal or nested-virtualisation-capable instance that costs $1.20 per hour and hosts 40 concurrent 4 GB-class sandboxes at your measured density. Compute alone is then $0.03 per sandbox-hour, far below the managed $0.36 per hour for the 4 GB tier.
But the instance must run when demand does not. If average utilisation is 25 percent because traffic is bursty and you hold headroom for spikes, the effective compute cost is four times higher, about $0.12 per used sandbox-hour. Add a warm-pool reserve of, say, 20 percent, and you are near $0.14. Now add engineering: if two engineers spend a quarter of their time on the platform and cost, in your market, an illustrative $20,000 per month combined, that is a fixed cost to amortise.
At the earlier 10,000 sessions of 12 minutes, the fleet uses 2,000 sandbox-hours per month. Managed at the 4 GB tier that is $720. Self-hosted compute at $0.14 per used hour is $280, but the engineering allocation dwarfs the saving by a factor of about 45. The crossover only arrives at scale: with the same illustrative per-hour rates, saving roughly $0.22 per sandbox-hour against $20,000 of monthly engineering cost needs on the order of 90,000 sandbox-hours a month, about 45 times this fleet, before pure cost favours building.
That crossover moves with every assumption, which is the point of showing it. Utilisation, density and the fixed-cost allocation swing the answer by an order of magnitude, so plug in your own numbers rather than anyone else’s.
Decision Matrix: Managed vs Partner vs Self-Hosted
There are really three options, not two. OpenAI-hosted sandboxes are the fully managed path. A partner provider, such as E2B, Daytona, Modal or Vercel, keeps the OpenAI harness while moving execution to a specialist. Self-hosted means your VPC or cluster and your isolation technology, connected through the bring-your-own option or a separate orchestrator.
| Criterion | OpenAI-hosted sandbox | Partner sandbox provider | Self-hosted Firecracker or gVisor |
|---|---|---|---|
| Time to first agent | Hours | Days | Weeks to months |
| Isolation transparency | Not detailed publicly | Provider documents its stack | Fully yours to choose and audit |
| Data residency control | Region and retention not specified | Depends on provider regions | Complete |
| Network egress policy | Limited to what OpenAI exposes | Provider features | Arbitrary |
| Marginal cost at low volume | Low | Low to moderate | High because of fixed cost |
| Marginal cost at high volume | Linear | Often lower than hosted | Lowest if utilisation is high |
| Operational burden | Minimal | Low | High, with on-call |
| Lock-in | High on harness and sandbox | Medium | Low on sandbox, depends on orchestrator |
| GPU or custom images | Memory, CPU and GPU options stated | Varies | Anything you can rack |
Read the matrix by your binding constraint. If you need an agent in front of users next month and your data is not regulated, the managed path wins on time and risk. If a regulator or customer contract demands a named isolation technology and a data-residency guarantee, the undisclosed items in OpenAI’s documentation push you toward a partner with a published stack or your own VPC.

Figure 4: A decision flow for the Agents API. Compliance and egress requirements gate the choice first, then volume and team capacity.
The bring-your-own path is the interesting middle
The announcement’s bring-your-own and VPC options deserve more attention than the headline price. They let a team adopt the Codex harness and the model loop while keeping execution inside its own network boundary, which resolves the most common enterprise objection: where does the generated code run and what can it touch?
The trade is complexity at the seam. You now operate an execution endpoint that must satisfy the harness’s interface, handle its file and command semantics, and authenticate its calls. Expect to spend real time on the failure paths, particularly timeouts, partial file writes and reconnection after a network blip mid-command. Pilot with a non-critical workflow first, and log every interface error, because beta interfaces change.
Where partners fit
A partner provider is often the best risk-adjusted choice for teams with moderate volume. You inherit someone else’s warm pools and isolation engineering, you get a documented security posture to hand to your auditors, and the per-hour rates reported by third parties sit well below the hosted 4 GB tier. The cost is a second vendor relationship and a second place where an outage can stall your agents.
One practical test: ask each candidate for p50 and p95 sandbox start latency under your burst pattern, not a quoted average. Start latency drives user-perceived responsiveness for interactive agents, and averages hide the tail that users actually feel.
Trade-offs, Gotchas, and What Goes Wrong
Idle billing is the silent multiplier. Agents spend a surprising share of wall-clock time waiting: on model calls, on human approval, on slow external APIs. If a sandbox bills while waiting, a session that does four minutes of work can bill thirty. Anthropic’s managed runtime is reported to meter only the time a session’s status is running, excluding idle time. OpenAI’s idle-time policy is not documented in the sources reviewed, so test it empirically: open a session, do nothing for ten minutes, and compare the invoice line.
The five-minute minimum punishes chatty designs. An agent that spawns a fresh sandbox per tool call, because that feels safer, pays the floor on every call. Reuse a session across a task and tear it down at the end, or you will pay for 25 minutes to do five one-minute operations.
Isolation you cannot inspect is a trust decision. The hosted sandboxes mirror OpenAI’s Codex and ChatGPT infrastructure, which is a reasonable signal of engineering investment, but it is not a technical description. If your threat model includes a malicious tenant or a model steered by injected instructions, you want a named boundary, a documented escape history, and the right to audit. Our guide to agentic AI security and prompt injection explains why the sandbox is the last line of defence once a model has been manipulated.
Egress is the real exfiltration path. A perfectly isolated sandbox with unrestricted outbound network access can still send your files to an attacker’s server in one HTTP request. Whatever runtime you choose, default to deny-all egress with an allowlist. Check what controls the managed option exposes before you put sensitive data in it.
Subagent fan-out can surprise you twice. It multiplies container spend and multiplies token spend, because each subagent carries its own context. A configured limit of three concurrent subagents is a ceiling on parallelism, not on cost. Set a per-task budget and a hard session timeout independent of the platform.
Beta churn. Parameter names, tier names and prices can change before general availability. Wrap the API, pin versions where the platform allows it, and keep a thin test suite that exercises a full agent run end to end, so a silent behaviour change shows up in CI rather than in production.
Concentration risk. Using OpenAI for the model, the harness and the sandbox puts three dependencies behind one vendor. That is efficient and also correlated: a single outage or policy change takes out all three. A partner sandbox or self-hosted execution plane breaks the correlation for one layer at the cost of integration work.
Anti-patterns to avoid
Do not treat the $0.03 figure as a budget; it is the smallest of four tiers for a 20-minute block. Do not let agents hold sessions open waiting for humans. Do not skip a kill-switch because the platform is managed. And do not assume that a managed runtime absolves you of data-handling obligations; the data still leaves your boundary unless you choose the VPC option, and your obligations travel with it.
Practical Recommendations
For most teams, the right answer is staged rather than binary. Start on the managed runtime to learn what your agents actually do: how long sessions run, how much memory they need, how often they fan out, and how much of the bill is tokens. Those measurements are worth more than any list price, and the managed path gets you them in days.
Then let the data choose. If tokens dominate and sessions are short, stay managed; the sandbox is noise. If sandbox spend is large and mostly CPU-bound with modest memory, price a partner provider against the hosted tier using your measured session shape. If compliance requires a named isolation boundary or residency guarantee, move execution into your VPC through the bring-your-own path, and treat full self-hosting as justified only when volume reaches the order of tens of thousands of sandbox-hours a month or when you need capabilities nobody sells.
Whichever path you take, design for exit. Keep your tools behind MCP or a thin function interface so they work with any harness, keep prompts and evals in your own repository, and record per-session cost, duration and idle time in your own telemetry. Our guide to agent evaluation harnesses and trajectory evals shows how to capture that data in a form you can replay against a different runtime.
Pilot checklist:
- Ask OpenAI in writing about idle billing, maximum session length, data retention, regions and subagent sandbox allocation.
- Run a representative task 50 times and record tokens, sandbox minutes billed, and wall-clock time.
- Test one idle session to confirm whether waiting time is billed.
- Set deny-by-default egress and a per-session timeout and budget before any real data.
- Wrap the API behind an adapter so a beta change is a one-file fix.
- Price one partner provider against the measured session shape.
- Re-run the cost model monthly while the product is in beta.
Frequently Asked Questions
What is the OpenAI Agents API?
It is a managed service, in public beta since 10 September 2026, for building and running cloud agents on OpenAI’s Codex harness. One call specifies task, model, tools and environment, and OpenAI runs the reasoning loop and can host the sandbox. It supports MCP, custom functions and web search, plus context compaction and subagents. OpenAI states there is no separate API fee beyond tokens and tools.
How much do OpenAI Agents API sandboxes cost?
According to a third-party summary of OpenAI’s pricing page, container sessions cost $0.03 per 20 minutes at 1 GB, $0.12 at 4 GB, $0.48 at 16 GB and $1.92 at 64 GB, billed by the minute with a five-minute minimum. Model tokens are billed separately. Rates can change during the beta, so check OpenAI’s pricing page before budgeting.
Can I use my own sandbox with the Agents API?
Yes. OpenAI’s announcement describes a bring-your-own option and first-class integrations with Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle and Runloop, plus Vercel. It also mentions deployments inside your own VPC. The integration details, such as the exact interface your endpoint must implement, are in OpenAI’s developer documentation, which you should read before committing.
Is a managed sandbox secure enough for untrusted code?
OpenAI says hosted sandboxes mirror its Codex and ChatGPT infrastructure but has not detailed the isolation technology in the sources reviewed. That makes it hard to judge against your threat model. For low-risk workloads it is probably adequate. For regulated data or hostile inputs, prefer a provider or a self-hosted runtime with a documented boundary such as Firecracker microVMs, and restrict network egress.
When does self-hosting an agent runtime make sense?
Self-hosting pays when volume is high enough to amortise platform engineering, when compliance demands a specific isolation technology or data residency, or when you need custom hardware. In the illustrative model above, the crossover sat near tens of thousands of sandbox-hours per month. Below that, managed or partner sandboxes usually cost less once you count engineering time and on-call.
Does the Agents API replace frameworks like LangGraph?
Not entirely. It replaces the orchestration and execution you would otherwise build for coding-style, tool-using agents. Frameworks still win when you need custom control flow, deterministic graphs, multi-vendor models or fine-grained state handling. Many teams will use the Agents API for open-ended task agents and a framework for structured workflows, keeping tools behind MCP so both can share them.
Further Reading
- AI agent frameworks benchmark: LangGraph, OpenAI and Google ADK for the orchestration side of the build-versus-buy choice.
- AI agent sandboxes compared: Firecracker, gVisor and Kata for isolation internals, density and startup trade-offs.
- OpenAI Astra and opaque recurrence explained for another recent OpenAI agent-oriented release.
- Agentic AI security and prompt injection for the threat model that makes sandboxing necessary.
- External: Introducing the Agents API (OpenAI) and the OpenAI developer community announcement.
- External: OpenAI Agents API Pricing: The Sandbox Is the Fee (tokencost.app), the third-party pricing summary cited above.
By Riju — about
