Repository navigation
Proposal: Split Supervisor and Agent into Separate Pods with gVisor Isolation #981
Description
Activity
Thanks for the detailed proposal. I am generally aligned with the direction, especially:
- Using Kubernetes image volumes to inject trusted OpenShell components.
- Treating the agent image as fully untrusted.
- Supporting gVisor through Kubernetes
RuntimeClass.
Some questions and comments:
- What scale or resource constraint requires a shared 1:N supervisor instead of a simpler 1:1 supervisor-to-sandbox topology? My expectation is that supervisor resource usage should be minimal, so the added complexity around sandbox identity, noisy-neighbor isolation, scaling behavior, and per-sandbox Kubernetes NetworkPolicy management, etc is something I'd like to avoid.
- Why should the supervisor run as a separate pod rather than a sidecar-style companion for the sandbox workload? If the supervisor is dedicated 1:1, the proposal should explain what isolation or operational benefit we get from a separate pod versus the extra service discovery, cert, lifecycle, and NetworkPolicy wiring.
- Configure gVisor/runtime selection through
SandboxProfile(see proposal below), not directly as a low-level network mode setting. Move sideloading, SSH, Landlock, and cert behavior intoSandboxProfileentity as well. - I have no objection to disabling SSH by default. I'm curious to learn more about what workloads don't require a TTY.
- The Falco/SSH concern may be outdated: we are no longer opening SSH ports directly; SSH traffic is multiplexed over a gRPC stream.
- I'd like to give "2.5 Supervisor CA Trust" a little more thought. Option C seems the best of the options, but want to consider some other alternatives or integrating it with components like cert manager.
- Changes to process identity binding aren't ideal. Would like to give this some more thought. I don't object to shipping an initial implementation without this and continuing to iterate on capabilities.
Net: I support the design and think we should proceed. I would scope the first version to a 1:1 sandbox/supervisor model,
SandboxProfile-driven configuration, trusted image-volume injection, untrusted agent images, and gVisorRuntimeClasssupport.SandboxProfile Proposal
SandboxProfileshould be an admin-managed application entity that describes how a sandbox is run and configured at the infrastructure levelSandboxProfilecan also define provider posture. This should build on the provider proposal in #896:How It Works
- An administrator creates one or more named
SandboxProfileobjects. CreateSandboxreferences a profile by name.- The gateway resolves the profile at sandbox creation time.
- The resolved profile is snapshotted onto the sandbox so later profile edits
do not silently mutate running sandboxes. - The compute driver renders the resolved profile into backend resources.
- The sandbox policy is applied separately by the supervisor/runtime.
- Attached providers contribute a separate provider-policy layer, as described
in issue Enhanced Provider Management #896, rather than being merged directly into the base policy.
For Kubernetes, the driver would use the resolved profile to create:
- Agent pod.
- Dedicated supervisor companion pod.
- Service or internal address for agent-to-supervisor traffic.
- NetworkPolicy scoped to that one sandbox.
- Image-volume mounts for trusted OpenShell components.
- Cert material according to the agreed sandbox cert strategy.
The first split-pod version should only support a dedicated 1:1 supervisor. Shared supervisors can be reconsidered later after identity, policy scoping, and noisy-neighbor behavior are designed.
Suggested Shape
metadata: name: gvisor-dedicated profile: driver: kubernetes topology: mode: sidecar supervisor: dedicated runtime: agentRuntimeClassName: gvisor runAsNonRoot: true allowPrivilegeEscalation: false capabilities: drop: ["ALL"] trustedInjection: strategy: imageVolume image: ghcr.io/nvidia/openshell-supervisor:latest enabledFeatures: - supervisorEntrypoint - landlock network: enforcement: kubernetesNetworkPolicy requireEnforcementVerification: true proxy: required: true port: 3128 dns: mode: review certs: strategy: sandboxDefault providers: allowedTypes: ["github", "claude", "openai", "nvidia"] requiredTypes: [] runtimeAttach: true composeProviderPolicy: true proxySideCredentialInjection: true inference: autoConfigure: true localRouting: true access: ssh: mode: review observability: ocsfLogSource: supervisor
Hi folks,
With regards to 2.5 "Supervisor CA Trust", I want to propose that we also consider an alternative option, which will not introduce dependency to
cert-manager. The implementation uses K8s primitives and is based on PodCertificateRequests and ClusterTrustBundles in which the gateway will also act as a custom signer. See also the following links on how pods request these volumes:- https://kubernetes.io/docs/concepts/storage/projected-volumes/#podcertificate
- https://kubernetes.io/docs/concepts/storage/projected-volumes/#clustertrustbundle
I am sharing my draft notes that can serve as inspiration for further enhancing this proposal. Disclaimer: there are probably some rough edges and things that can be considered optional or require further discussion. Let me know what you think.
Details
Alternative for Supervisor CA Trust: Kubernetes CSR API + ClusterTrustBundle
Re: Section 2.5 "Supervisor CA Trust" — Options A/B/C
Proposal: Option D — Gateway as CSR Signer + Pod Certificate Volumes + ClusterTrustBundle
Instead of each supervisor generating its own ephemeral CA and manually distributing trust, we propose using Kubernetes-native pod certificate provisioning:
- Gateway acts as a custom CSR signer (
openshell.ai/proxy) and publishes its CA via ClusterTrustBundle - Supervisor pods receive their MITM signing cert automatically via
podCertificateprojected volumes — kubelet handles key generation, CSR submission, signing, and rotation - Agent pods contain a sidecar proxy container that receives a client identity cert via
podCertificate, and trust the MITM CA viaclusterTrustBundleprojected volumes. The untrusted agent container has no access to cert keys.
No manual cert distribution. No init containers. No Secrets for cert delivery. No code in supervisor or agent to interact with the K8s API.
Why the gateway as signer?
The gateway already owns the cluster PKI today:
openshell-bootstrapgenerates the root CA at cluster init (crates/openshell-bootstrap/src/pki.rs)- The gateway stores it in K8s Secrets (
openshell-server-tls,openshell-client-tls) - Supervisors already authenticate to the gateway via mTLS using certs issued by this CA
- The gateway has a stable lifecycle (StatefulSet) — supervisors are ephemeral per-sandbox
A signer-per-supervisor is problematic because:
- Cluster-scoped resource pollution — each supervisor creates/destroys a ClusterTrustBundle
- Startup ordering — agent pod can't start until its specific supervisor publishes its CA
- 1:N topology breaks — shared supervisor model requires a single CA anyway
- No revocation authority — supervisor can't revoke its own CA if compromised; gateway can
Architecture
┌───────────────────────────────────────────────────────────────────────────┐ │ GATEWAY (openshell-server) │ │ │ │ Already has: New responsibilities: │ │ • Bootstrap CA keypair • MITM CA keypair (separate from │ │ • K8s Secret management bootstrap CA for defense-in-depth)│ │ • mTLS acceptor for supervisors • CSR signer controller │ │ • Sandbox lifecycle management (signerName: openshell.ai/proxy) │ │ • ClusterTrustBundle publisher │ │ │ │ On startup: │ │ 1. Generate (or load) MITM proxy CA │ │ 2. Publish CA cert → ClusterTrustBundle "openshell-proxy-ca" │ │ 3. Watch PodCertificateRequests for signerName "openshell.ai/proxy" │ │ 4. Sign supervisor certs as intermediates (CA:TRUE, pathLen:0) │ │ 5. Sign agent certs as client leaves (client auth usage) │ └───────────────────────────────────┬───────────────────────────────────────┘ │ ┌─────────────────────┼─────────────────────┐ │ │ │ podCertificate │ podCertificate + │ projected volume │ clusterTrustBundle │ (intermediate CA cert) │ projected volumes ▼ ▼ ┌─────────────────────────────────┐ ┌──────────────────────────────────────┐ │ SUPERVISOR POD │ │ AGENT POD │ │ │ │ │ │ podCertificate: │ │ Sidecar container (trusted): │ │ signerName: │ │ podCertificate: │ │ openshell.ai/proxy │ │ signerName: openshell.ai/proxy │ │ → intermediate CA cert │ │ → client cert CN=agent-<pod-uid> │ │ → can sign leaf certs │ │ Runs local proxy on localhost:3128 │ │ │ │ mTLS to supervisor proxy │ │ On HTTPS CONNECT: │ │ │ │ 1. Generate leaf cert for host │ │ clusterTrustBundle: │ │ 2. Present leaf to agent │ │ signerName: openshell.ai/proxy │ │ 3. Agent trusts it (chain → │ │ → gateway MITM CA (public) │ │ intermediate → gateway CA) │ │ │ │ 4. Decrypt + inspect traffic │ │ Agent container (untrusted): │ │ 5. Re-encrypt to upstream │ │ HTTP_PROXY=http://localhost:3128 │ │ │ │ No certs, no mTLS config │ │ No CA generation! │ │ Trusts gateway CA for MITM only │ │ No K8s API access! │ │ │ └─────────────────────────────────┘ └──────────────────────────────────────┘How the signer decides what cert to issue
The
podCertificateprojected volume does not let the requester specify "I want a CA cert" or "I want a client leaf." The volume source only declares thesignerName,keyType, andmaxExpirationSeconds. The signer decides what kind of certificate to issue based on the pod's identity.When kubelet creates a
PodCertificateRequest, it includes full pod metadata: namespace, name, UID, ServiceAccount, and node. The gateway signer controller looks up the pod's labels via the K8s API and issues the appropriate cert profile:// Gateway signer logic example (crates/openshell-server/src/cert_signer.rs) let pod = k8s_client.get_pod(&request.namespace, &request.pod_name).await?; // 1. Validate ownerReferences — pod must be owned by the OpenShell driver let valid_owner = pod.owner_references.iter().any(|ref| { ref.kind == "Job" || ref.kind == "Pod" // driver-created pods && ref.controller == Some(true) // For split-pod: could also check that owner is a specific StatefulSet/ReplicaSet // managed by the openshell-driver-kubernetes controller }); if !valid_owner { return deny_request(&request, "pod not owned by OpenShell driver"); } // 2. Validate ServiceAccount — must be the expected SA for the role let expected_sa = match pod.labels.get("openshell.ai/role").map(|s| s.as_str()) { Some("supervisor") => "openshell-supervisor", Some("agent") => "openshell-agent", _ => return deny_request(&request, "unknown role"), }; if pod.service_account.as_deref() != Some(expected_sa) { return deny_request(&request, "unexpected ServiceAccount for role"); } // 3. Validate namespace — must be an OpenShell-managed namespace if !is_managed_namespace(&pod.namespace) { return deny_request(&request, "pod not in managed namespace"); } // 4. Decide cert profile based on validated role label match pod.labels.get("openshell.ai/role").map(|s| s.as_str()) { Some("supervisor") => { // Issue intermediate CA cert: // basicConstraints: CA:TRUE, pathLenConstraint:0 // keyUsage: keyCertSign, cRLSign // nameConstraints: excluded subtree (see below) // CN: supervisor-<sandbox-id> sign_as_intermediate_ca(&request, &mitm_ca_key) } Some("agent") => { // Issue client leaf cert: // basicConstraints: CA:FALSE // extKeyUsage: clientAuth // CN: agent-<pod-uid> sign_as_client_leaf(&request, &mitm_ca_key) } _ => unreachable!() // already checked above }
Signer validation guarantees:
The signer does NOT trust labels alone. Before issuing any certificate, it validates the full chain of ownership:
ownerReferences— pod must be created by the OpenShell driver (controller=true). A rogue pod created viakubectl runor by another controller will be rejected even if it has the correct labels.- ServiceAccount — must match the expected SA for the claimed role. Prevents a pod with
role=supervisorlabel but the wrong SA from getting an intermediate CA. - Namespace — must be an OpenShell-managed namespace. Prevents cross-namespace attacks.
- Labels — determine cert profile only after ownership is validated.
This defense-in-depth ensures that even if an attacker can
patchpod labels (requires RBAC), they still can't get an intermediate CA cert because:- The pod's ownerReferences are immutable after creation (set by the controller that created it)
- The pod's ServiceAccount is immutable after creation
- Both must match the expected values for a supervisor
Name Constraints on the supervisor intermediate cert
The supervisor intermediate needs to sign leaf certs for arbitrary external hostnames (whatever the agent connects to). This means we can't use a permitted subtree (allowlist) — the supervisor doesn't know in advance which hosts the agent will call.
Instead, we use an excluded subtree to prevent the supervisor from signing certs that could impersonate internal cluster services:
X509v3 Name Constraints: critical Excluded: DNS: cluster.local # all cluster-internal services DNS: svc.cluster.local # all Kubernetes services DNS: openshell.svc.cluster.local # gateway and other OpenShell services DNS: localhost # prevent localhost impersonation IP: 10.0.0.0/8 # cluster pod/service CIDRs IP: 172.16.0.0/12 IP: 192.168.0.0/16 IP: 127.0.0.0/8 # loopbackThis means a compromised supervisor intermediate:
- CAN sign leaf certs for
api.openai.com,github.com, etc. (needed for MITM) - CANNOT sign certs for
openshell.openshell.svc.cluster.local(the gateway) - CANNOT sign certs for any internal cluster service
- CANNOT sign certs for localhost or private IP ranges
Why excluded (not permitted) subtree:
- The supervisor must sign for arbitrary external domains — a permitted subtree can't enumerate the internet
- The actual threat is impersonating internal services (gateway, other supervisors, K8s API)
- Excluding
cluster.local+ private IPs covers all internal attack surfaces - X.509 validators enforce Name Constraints automatically — a leaf cert for
openshell.svc.cluster.localsigned by this intermediate will be rejected by any standards-compliant TLS library
Open question: What should the excluded subtree contain?
Agents may legitimately need to reach internal cluster services (e.g., inference backends at
vllm.mynamespace.svc.cluster.local, vector databases, tool APIs). A blanket exclusion ofcluster.localwould break these use cases. Options:Option Excluded subtrees Trade-off A: Exclude only OpenShell + kube-system openshell.svc.cluster.local,kube-system.svc.cluster.local,localhost,127.0.0.0/8Protects control plane. Allows agents to reach other internal services via MITM. B: Exclude all cluster.local cluster.local,svc.cluster.local,localhost, all private IPsMaximum protection. Agents cannot reach internal services through the MITM proxy (would need a passthrough/bypass rule). C: Configurable per deployment Operator-defined via ConfigMap Most flexible. Defaults to Option A. Operators can tighten to Option B if agents don't need internal access. This decision depends on the expected agent workloads and whether internal service access should go through MITM inspection or be handled differently (e.g., direct NetworkPolicy-allowed connections that bypass the proxy).
Alternative: Decide based on
userAnnotationsinstead of labelsThe
podCertificatespec supports auserAnnotationsmap — arbitrary key-value pairs passed through to the signer in thePodCertificateRequest. The driver could encode the desired cert profile directly in the volume spec:# Supervisor pod — requests intermediate CA via annotation - podCertificate: signerName: "openshell.ai/proxy" keyType: ECDSAP256 maxExpirationSeconds: 14400 keyPath: tls.key certificateChainPath: tls.crt userAnnotations: openshell.ai/cert-profile: "intermediate-ca" openshell.ai/sandbox-id: "<sandbox-id>" # Agent pod — requests client leaf via annotation - podCertificate: signerName: "openshell.ai/proxy" keyType: ECDSAP256 maxExpirationSeconds: 14400 keyPath: client.key certificateChainPath: client.crt userAnnotations: openshell.ai/cert-profile: "client-leaf" openshell.ai/sandbox-id: "<sandbox-id>"
Gateway signer logic with
userAnnotations:let profile = request.user_annotations.get("openshell.ai/cert-profile"); match profile.map(|s| s.as_str()) { Some("intermediate-ca") => { // Validate: pod actually has label openshell.ai/role=supervisor // (annotations alone are not trusted — still cross-check labels) sign_as_intermediate_ca(&request, &mitm_ca_key) } Some("client-leaf") => { sign_as_client_leaf(&request, &mitm_ca_key) } _ => deny_request(&request, "missing or unknown cert-profile annotation") }
Trade-offs:
Label-based (recommended) userAnnotations-based Simplicity Signer only needs K8s API to read labels Signer reads annotations from request directly (no extra API call) Security Labels are set by driver, validated server-side Annotations are in the pod spec — same trust boundary (driver sets them) Flexibility Fixed: one role = one cert type Extensible: can add new profiles without new labels Auditability Role visible in kubectl get pods --show-labelsProfile only visible in PodCertificateRequest object Extra data Requires separate lookup for sandbox-id Can pass sandbox-id directly in annotations Recommendation: Use labels as the primary decision mechanism (simpler, already the pattern in the codebase), but accept
userAnnotationsas supplementary metadata (e.g.,sandbox-idfor binding validation). The signer should always cross-check annotations against pod labels — never trust annotations alone, since the pod spec is the source of truth for role.Security property (both approaches): An untrusted workload cannot request CA signing privileges. Whether decided by label or annotation, both are set by the driver when constructing the pod spec — the workload running inside the pod has no ability to modify either after creation. The signer additionally validates:
- ownerReferences (immutable, proves the pod was created by the OpenShell driver)
- ServiceAccount (immutable, must match expected SA for the claimed role)
- Namespace (must be OpenShell-managed)
Even if an attacker gains
patch podsRBAC and relabels a pod torole=supervisor, the signer rejects it because ownerReferences and ServiceAccount are immutable after pod creation and won't match.Both the supervisor and agent pod specs look similar from the volume perspective (differing only in annotations if used):
# Supervisor pod — gets intermediate CA (signer decides based on role label) - podCertificate: signerName: "openshell.ai/proxy" keyType: ECDSAP256 maxExpirationSeconds: 14400 keyPath: tls.key certificateChainPath: tls.crt # Agent pod — gets client leaf (signer decides based on role label) - podCertificate: signerName: "openshell.ai/proxy" keyType: ECDSAP256 maxExpirationSeconds: 14400 keyPath: client.key certificateChainPath: client.crt
Same
signerName, samekeyType— but the resulting cert is fundamentally different because the signer makes the decision based on who is asking, not what they ask for.How MITM decryption works (end-to-end)
This is the critical flow that makes L7 inspection possible in the split-pod model:
Agent process Proxy sidecar Supervisor proxy Upstream (e.g. api.openai.com) │ │ │ │ │ CONNECT │ │ │ │ api.openai.com │ │ │ ├───────────────────►│ │ │ │ (plaintext HTTP │ CONNECT │ │ │ to localhost) │ api.openai.com:443 │ │ │ ├─────────────────────►│ │ │ │ (mTLS client cert │ │ │ │ proves agent ID) │ │ │ │ │ │ │ │ HTTP 200 │ │ │ HTTP 200 │◄─────────────────────┤ │ │◄───────────────────┤ │ │ │ │ │ │ │ TLS ClientHello │ │ │ ├───────────────────►│ (passthrough) │ │ │ ├─────────────────────►│ │ │ │ │ TLS ClientHello │ │ │ ├─────────────────────────────►│ │ │ │ TLS ServerHello + real cert │ │ │ │◄─────────────────────────────┤ │ │ │ │ │ TLS ServerHello + │ │ (Supervisor has real TLS │ │ MITM leaf cert │ (passthrough) │ session to upstream) │ │◄───────────────────┤◄─────────────────────┤ │ │ │ │ │ │ (Agent validates: │ │ │ │ leaf signed by │ │ │ │ supervisor int., │ │ │ │ int. signed by │ │ │ │ gateway CA in │ │ │ │ trust store) │ │ │ │ │ │ │ │ Application data │ │ │ ├───────────────────►├─────────────────────►│ │ │ │ │ ┌──────────────────────┐ │ │ │ │ │ Decrypt with MITM │ │ │ │ │ │ key → plaintext HTTP │ │ │ │ │ │ │ │ │ │ │ │ OPA policy check │ │ │ │ │ │ Credential injection │ │ │ │ │ │ OCSF logging │ │ │ │ │ │ Inference routing │ │ │ │ │ │ │ │ │ │ │ │ Re-encrypt with │ │ │ │ │ │ upstream TLS session │ │ │ │ │ └──────────────────────┘ │ │ │ │ │ │ │ │ Application data (HTTP) │ │ │ ├─────────────────────────────►│ │ │ │ │Certificate trust chain:
Gateway MITM CA (in agent's trust store via ClusterTrustBundle) └── Supervisor intermediate cert (issued via podCertificate, pathLen:0) └── Per-hostname leaf cert (generated at runtime by supervisor) e.g. CN=api.openai.com, SAN=api.openai.comWhy the agent trusts the MITM cert:
- Agent's trust store contains the gateway MITM CA (via
clusterTrustBundleprojected volume →SSL_CERT_FILE) - Supervisor presents a leaf cert for
api.openai.comsigned by its intermediate - Intermediate is signed by the gateway MITM CA
- Standard X.509 chain validation succeeds → agent accepts the connection
- Supervisor can now read plaintext HTTP, apply policy, inject credentials, then re-encrypt to upstream
Why the supervisor can decrypt:
- It holds the private key for the leaf cert it just generated (in-memory, per-connection)
- The agent encrypted the TLS session using the leaf cert's public key
- Supervisor decrypts with the corresponding private key → plaintext HTTP
How the agent gets its client certificate
The agent pod uses a sidecar proxy container to handle mTLS authentication to the supervisor. The untrusted agent container never has access to cert material:
Agent pod architecture: ┌───────────────────────────────────────────────────────────────────┐ │ │ │ ┌──────────────────────────────┐ ┌────────────────────────────┐ │ │ │ Sidecar container (trusted) │ │ Agent container │ │ │ │ Image: openshell/proxy-client│ │ Image: untrusted │ │ │ │ │ │ │ │ │ │ Mounts: │ │ Mounts: │ │ │ │ - podCertificate volume │ │ - clusterTrustBundle only │ │ │ │ (client.key + client.crt) │ │ (gateway CA, public) │ │ │ │ │ │ │ │ │ │ Runs: │ │ Env: │ │ │ │ - Local proxy on │ │ - HTTP_PROXY= │ │ │ │ localhost:3128 │ │ http://localhost:3128 │ │ │ │ - Presents client cert on │ │ - SSL_CERT_FILE= │ │ │ │ every connection to │ │ /etc/ssl/certs/ │ │ │ │ supervisor │ │ proxy-ca.pem │ │ │ │ │ │ │ │ │ │ The sidecar: │ │ The agent: │ │ │ │ ✓ Holds cert key │ │ ✗ No cert key access │ │ │ │ ✓ Authenticates to supervisor│ │ ✗ No mTLS config │ │ │ │ ✓ Small, auditable image │ │ ✓ Trusts MITM CA (for L7) │ │ │ │ ✓ Runs as separate process │ │ ✓ Uses proxy transparently │ │ │ └──────────────────────────────┘ └────────────────────────────┘ │ │ ▲ │ │ │ │ localhost:3128 (plaintext) │ │ │ └───────────────────────────────┘ │ └───────────────────────────────────────────────────────────────────┘Flow:
- kubelet generates keypair for sidecar container's
podCertificatevolume - kubelet creates PodCertificateRequest → gateway signs as client leaf
- kubelet mounts cert into sidecar container only (not agent container)
- Sidecar starts local proxy on
localhost:3128 - Agent container starts with
HTTP_PROXY=http://localhost:3128 - Agent makes HTTPS requests → routed to local sidecar → sidecar forwards to supervisor with mTLS
Security property: The agent container cannot access the client cert key because:
- The
podCertificatevolume is mounted only in the sidecar container - Containers in the same pod have separate filesystem namespaces (unless sharing volumes explicitly)
- No volume mount path for the cert exists in the agent container spec
- Even with a container escape, the key lives in the sidecar's mount namespace
This mirrors the current single-pod architecture where the supervisor process holds TLS material and the agent child process never touches it — expressed as containers rather than processes.
Security consideration: Agent access to its own client cert key
The agent process must have access to the client cert and key because standard HTTP libraries handle
HTTPS_PROXYconnections directly — the agent's HTTP library opens the TCP+TLS connection to the supervisor proxy and presents the client cert during the TLS handshake. A sideloaded binary that exec-s into the agent cannot interpose on this.Two approaches:
Approach A: Local proxy sidecar (strongest isolation)
A separate sidecar container runs a local proxy that terminates the mTLS connection to the supervisor:
┌──────────────────────────────────────────────────────────────┐ │ Agent Pod │ │ │ │ ┌──────────────────────┐ ┌────────────────────────────┐ │ │ │ Agent container │ │ Sidecar proxy container │ │ │ │ │ │ │ │ │ │ HTTP_PROXY= │ TCP │ Holds client cert + key │ │ │ │ http://localhost: ├─────►│ mTLS to supervisor:3128 │ │ │ │ 3128 │ │ │ │ │ │ │ │ No cert volume mounted │ │ │ │ No certs, no TLS │ │ in agent container │ │ │ └──────────────────────┘ └────────────────────────────┘ │ └──────────────────────────────────────────────────────────────┘- Agent container has
HTTP_PROXY=http://localhost:3128(plaintext to sidecar, no certs) - Sidecar container mounts the
podCertificatevolume (client cert + key) - Sidecar authenticates to supervisor via mTLS
- Agent container has no access to cert material — volume is not mounted in its container
- The sidecar is a trusted OpenShell image (small, auditable)
This provides true isolation: the agent literally cannot read the cert key because it's in a different container's filesystem namespace.
Approach B: Accept agent access to its own client cert (simpler, sufficient)
The agent process has access to the client cert. This is acceptable because:
- The cert is a leaf —
CA:FALSE, cannot sign other certs - Scoped to one pod UID — CN=
agent-<pod-uid>, useless for impersonating other agents - Short-lived — expires with the sandbox (max 4h, auto-rotated), no long-term value
- Only useful for its intended purpose — authenticating to the supervisor proxy, which it's already supposed to do
- Supervisor still enforces policy — OPA evaluates every request regardless of auth
The worst an agent can do with its own client cert is... make requests to its own supervisor proxy, which is the designed behavior. The cert provides identity (the supervisor logs which agent made each request), not authorization (OPA decides what's allowed).
Recommendation: Use Approach A (sidecar) for maximum isolation. Fall back to Approach B if the pod spec complexity of a sidecar is unacceptable for the deployment model.
Approach A means the agent pod has two containers:
- Sidecar: trusted OpenShell image, mounts cert volume, runs local proxy
- Agent: untrusted image, no cert access, uses
HTTP_PROXY=http://localhost:3128
This mirrors the current single-pod architecture where the supervisor process holds TLS material and the agent child process never touches it — just deployed as containers instead of processes.
Kubernetes manifests
Supervisor pod spec:
volumes: - name: proxy-cert projected: sources: - podCertificate: signerName: "openshell.ai/proxy" keyType: ECDSAP256 maxExpirationSeconds: 14400 # 4 hours — short-lived, auto-rotated keyPath: tls.key certificateChainPath: tls.crt containers: - name: supervisor env: - name: OPENSHELL_PROXY_CA_CERT value: /etc/openshell-tls/proxy/tls.crt - name: OPENSHELL_PROXY_CA_KEY value: /etc/openshell-tls/proxy/tls.key volumeMounts: - name: proxy-cert mountPath: /etc/openshell-tls/proxy readOnly: true
Agent pod spec (sidecar handles mTLS, agent has no cert access):
volumes: - name: client-cert projected: sources: - podCertificate: signerName: "openshell.ai/proxy" keyType: ECDSAP256 maxExpirationSeconds: 14400 # 4 hours — short-lived, auto-rotated keyPath: client.key certificateChainPath: client.crt - name: proxy-ca projected: sources: - clusterTrustBundle: signerName: "openshell.ai/proxy" path: proxy-ca.pem containers: - name: proxy-sidecar image: openshell/proxy-client:latest volumeMounts: - name: client-cert mountPath: /etc/openshell-tls/client readOnly: true - name: proxy-ca mountPath: /etc/ssl/certs readOnly: true ports: - containerPort: 3128 name: proxy protocol: TCP - name: agent image: <untrusted-agent-image> env: - name: HTTP_PROXY value: http://localhost:3128 - name: HTTPS_PROXY value: http://localhost:3128 - name: SSL_CERT_FILE value: /etc/ssl/certs/proxy-ca.pem - name: REQUESTS_CA_BUNDLE value: /etc/ssl/certs/proxy-ca.pem - name: NODE_EXTRA_CA_CERTS value: /etc/ssl/certs/proxy-ca.pem volumeMounts: - name: proxy-ca mountPath: /etc/ssl/certs readOnly: true # NOTE: client-cert volume is NOT mounted in agent container
ClusterTrustBundle (published by gateway):
apiVersion: certificates.k8s.io/v1beta1 kind: ClusterTrustBundle metadata: name: openshell-proxy-ca spec: signerName: openshell.ai/proxy trustBundle: | -----BEGIN CERTIFICATE----- <gateway's MITM proxy CA cert> -----END CERTIFICATE-----
Gateway signer RBAC:
apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: openshell-csr-signer rules: - apiGroups: ["certificates.k8s.io"] resources: ["podcertificaterequests"] verbs: ["get", "list", "watch"] - apiGroups: ["certificates.k8s.io"] resources: ["podcertificaterequests/status"] verbs: ["update"] - apiGroups: ["certificates.k8s.io"] resources: ["signers"] resourceNames: ["openshell.ai/proxy"] verbs: ["sign"] - apiGroups: ["certificates.k8s.io"] resources: ["clustertrustbundles"] verbs: ["create", "update", "get"]
Agent-supervisor binding
Since all supervisors hold certs signed by the same gateway CA, an agent trusting the CA could connect through any supervisor. This mirrors today: in the single-pod model, binding is enforced by network namespace isolation.
In the split-pod model, NetworkPolicy provides binding (L3/L4), and mTLS client cert provides identity:
# NetworkPolicy: agent can only reach its assigned supervisor egress: - to: - podSelector: matchLabels: openshell.ai/sandbox-id: <this-sandbox-id> ports: - port: 3128
Even if an agent somehow bypassed the NetworkPolicy, the supervisor validates the client cert's CN (
agent-<pod-uid>) against its expected agent — rejecting connections from unknown agents.Comparison with RFC Options A/B/C
A (init container) B (Secret) C (env var) D (Pod Cert + CTB) Startup ordering Agent waits for supervisor Agent waits for Secret Agent waits for supervisor Both start independently CA rotation Restart agent pods Update Secret + restart Restart agent pods Auto-refresh by kubelet 1:N topology Per-supervisor init Per-supervisor Secret Per-supervisor config Single trust bundle Agent identity Separate mechanism Separate mechanism Separate mechanism Built-in mTLS Security model Leaking cert = benign Leaking Secret = bad Leaking cert = benign Trust bundle = public Private key exposure In memory In K8s Secret In memory Never leaves node Supervisor K8s API Not needed Not needed Not needed Not needed Agent K8s API Not needed Not needed Not needed Not needed Cert rotation Manual Manual Manual Automatic (kubelet) Agent impersonation No protection No protection No protection Impossible (mTLS) Limitations and mitigations
Limitation Mitigation podCertificate is beta (K8s 1.35) Feature gate PodCertificateRequest. For older clusters, driver submits CSR manually + mounts result as Secret. Detect API at runtime.ClusterTrustBundle is beta (K8s 1.33) Feature gate ClusterTrustBundleProjection. Fallback: mount CA cert as Secret volume in agent pods.Gateway signer controller (~150 lines) Watch PodCertificateRequests, validate pod identity + labels, sign with MITM CA, update status. Gateway already has K8s client. Supervisor cert is an intermediate CA Issued with basicConstraints: CA:TRUE, pathLenConstraint:0and Name Constraints (see open question below on exact exclusion scope).rcgen(already in workspace) supports both.Pod startup blocked until cert ready By design — kubelet won't start pod until signer responds. Gateway signer responds in seconds (same cluster). Agent cert key protection Sidecar container holds the cert; volume not mounted in agent container. Filesystem namespace isolation between containers — agent cannot access cert material. Sidecar adds a container to agent pod Small, purpose-built OpenShell image (~10MB). Same pattern as Istio/Envoy sidecars. Can be injected via mutating webhook for transparency. Implementation sketch
// In crates/openshell-server/src/cert_signer.rs (new ~150 line module) // Gateway watches PodCertificateRequests with signerName "openshell.ai/proxy" // On new request: // 1. Extract pod identity (namespace, name, UID) from request // 2. Look up pod via K8s API, validate: // - ownerReferences: must be controller-owned by OpenShell driver // - ServiceAccount: must match expected SA for role // - Namespace: must be OpenShell-managed // 3. Decide cert profile based on openshell.ai/role label: // - "supervisor" → intermediate CA (CA:TRUE, pathLen:0, certSign) // - "agent" → client leaf (CA:FALSE, clientAuth, CN=agent-<pod-uid>) // - anything else → deny // 4. Sign with MITM CA, update request status with certificate chain // In crates/openshell-sandbox/src/l7/tls.rs (modify existing) // Add SandboxCa::from_files(cert_path, key_path) -> Result<Self> // Startup: try from_files(), fall back to generate() for backwards compat // In crates/openshell-sandbox/src/proxy.rs (modify existing) // Add mTLS client verification on accepted connections: // - Require client cert on CONNECT // - Validate cert chain against gateway CA // - Extract agent identity from cert CN // - Use identity for per-agent policy evaluation + OCSF logging // In crates/openshell-driver-kubernetes/src/driver.rs (modify existing) // Supervisor pod spec: add podCertificate projected volume // Agent pod spec: // - Add sidecar container (openshell/proxy-client image) // - Mount podCertificate volume into sidecar only // - Mount clusterTrustBundle volume into both containers // - Set HTTP_PROXY=http://localhost:3128 in agent container // Fallback: if cluster doesn't support podCertificate, use Secret-based approach
Certificate lifetimes and CA rotation
Short-lived certificates instead of revocation infrastructure
Since
podCertificateprovides automatic rotation (kubelet renews before expiry with jitter), there's no operational cost to short cert lifetimes. We recommend:Certificate Lifetime Rationale Supervisor intermediate 2-4 hours Limits blast radius of stolen intermediate. Sandbox lifetime is typically hours, not days. kubelet refreshes transparently. Agent client cert 2-4 hours Same reasoning. Pod lifetime is the upper bound anyway. Gateway MITM CA 90 days Root CA rotation is disruptive, so longer lifetime. Protected by K8s Secret RBAC + gateway-only access. With 2-4 hour intermediate lifetimes:
- A stolen intermediate key is useless after a few hours (no CRL/OCSP needed)
- A compromised supervisor that's killed has its cert expire quickly — no long-lived credentials to chase
- kubelet handles rotation automatically with jitter to avoid thundering herd
# Supervisor pod — 4 hour cert lifetime - podCertificate: signerName: "openshell.ai/proxy" keyType: ECDSAP256 maxExpirationSeconds: 14400 # 4 hours keyPath: tls.key certificateChainPath: tls.crt
Graceful MITM CA rotation (dual-trust transition)
When the gateway MITM CA must be rotated (planned rotation, compromise, key ceremony), a zero-downtime strategy:
Timeline: ───────────────────────────────────────────────────────────────────── T+0: Gateway generates new MITM CA (CA-new) ClusterTrustBundle updated to contain BOTH CA-old + CA-new: trustBundle: | -----BEGIN CERTIFICATE----- <CA-old cert> -----END CERTIFICATE----- -----BEGIN CERTIFICATE----- <CA-new cert> -----END CERTIFICATE----- Agents now trust both CAs (projected volume refreshes ~60s). Existing supervisors still hold intermediates signed by CA-old — still valid. T+1h: Gateway signer switches to signing NEW intermediates with CA-new. As supervisor certs rotate (every 2-4h), they naturally get CA-new intermediates. Agents trust both → no disruption during transition. T+8h: All supervisor intermediates have rotated to CA-new. (Worst case: max cert lifetime of 4h means all have rotated within 4h + jitter) T+9h: Gateway removes CA-old from ClusterTrustBundle: trustBundle: | -----BEGIN CERTIFICATE----- <CA-new cert only> -----END CERTIFICATE----- Agents drop trust for CA-old on next projected volume sync. Any remaining CA-old intermediates (shouldn't exist) stop working. ─────────────────────────────────────────────────────────────────────Key properties:
- Zero downtime — agents trust both CAs during the transition window
- Self-healing — short-lived supervisor certs naturally rotate to the new CA without intervention
- No pod restarts needed — ClusterTrustBundle projected volume updates in-place
- Safe rollback — if something goes wrong, re-add CA-old to the trust bundle
Emergency rotation (CA compromise):
- Same flow but compressed: immediately add CA-new, stop signing with CA-old
- Accept a brief window (up to max cert lifetime) where compromised intermediates still work
- This is why short cert lifetimes matter — 4h max exposure vs. 24h
Summary
Fully Kubernetes-native certificate provisioning with zero custom cert management code:
- Gateway holds MITM CA + runs signer controller for
openshell.ai/proxy - Supervisor pods get intermediate CA cert via
podCertificate— kubelet handles everything - Supervisor generates per-hostname leaf certs at runtime (same
CertCacheas today) - Agent pods have a sidecar proxy container that gets client cert via
podCertificate - Sidecar handles mTLS to supervisor — agent container has no cert access
- Agent pods trust the MITM CA via
clusterTrustBundleprojected volume (for TLS validation of MITM certs) - Supervisor validates sidecar's client cert on every CONNECT — cryptographic identity
- NetworkPolicy provides L3/L4 binding; mTLS provides identity
No Secrets for cert delivery. No init containers. No K8s API access from workloads. No manual rotation. Private keys never leave the node. Agent container never has access to client cert key material.
References
Kubernetes documentation:
- Certificate Signing Requests — CSR API, custom signers, ClusterTrustBundles
- Projected Volumes: podCertificate — automatic pod certificate provisioning via kubelet
- Projected Volumes: clusterTrustBundle — trust anchor distribution to pods
Kubernetes Enhancement Proposals (KEPs):
- KEP-4317: Pod Certificates — podCertificate projected volume source (beta in K8s 1.35)
- KEP-3257: ClusterTrustBundles — ClusterTrustBundle resource and projected volume source (beta in K8s 1.33)
- KEP-1513: Certificate Signing Requests — CSR API v1 (GA since K8s 1.19)
Reacted by Konstantinos AngelopoulosPer-binary destination restrictions. In the current model, the proxy resolves which binary opened a network connection by reading /proc//exe from the supervisor (which shares a PID namespace with the agent). In the split-pod model, the supervisor and agent are in different pods with different PID namespaces. The supervisor proxy sees TCP connections from the agent pod's IP but cannot determine which binary initiated the connection.
Yes but I'll be blunt: the "which binary initiated the connection" is broken anyways today by
LD_PRELOAD.Option C (cooperative agent image): The agent image is expected to trust the supervisor CA via REQUESTS_CA_BUNDLE or SSL_CERT_FILE environment variables pointing to a mounted cert. No image modification needed.
Yeah this is kind of a best practice today, but IMO the real fix is still podman-container-tools/podman#2317
Hi @drew
thanks for your general "thumbs up". Let me provide some (yet partial) feedback:
What scale or resource constraint requires a shared 1:N supervisor instead of a simpler 1:1 supervisor-to-sandbox topology?
We do expect a massive number of sandboxed agents per Kubernetes cluster (multiple 10k, possibly exceeding 100k). At that scale the overhead will be noticeable. We also expect a hight number of agents to be in a warm state.
so the added complexity around sandbox identity,
Agreed, that does increase complexity, however that should be addressable by Dimitar's proposal above.
noisy-neighbor isolation
I am not sure how this would be different in the shared model. We would of course have replicas of the supervisors to avoid bottlenecks.
per-sandbox Kubernetes NetworkPolicy management
That should not be required, one single network policy of this kind should be sufficient if we share supervisors:
apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: name: sandbox-to-supervisor namespace: supervisor-namespace # Apply this in the namespace where supervisors live spec: podSelector: matchLabels: type: supervisor # Targets the supervisor pods policyTypes: - Ingress ingress: - from: - namespaceSelector: {} # Match ALL namespaces (or restrict with matchLabels) podSelector: matchLabels: type: sandbox # Only pods labeled type=sandboxWhy should the supervisor run as a separate pod rather than a sidecar-style companion for the sandbox workload
That is indeed interesting. I can see that it may invalidate the 1:1 supervisor topology as it appears to be much simpler compared to "remote pod" model (it also avoids network policies for every pair). There is however a little caveat: will this model also work with gVisor? Information on this is somewhat inconsistent.
For this scenario we discussed an istio like approach, e.g. forward all network traffic to the supervisor process, make it impossible for the agent to communicate directly with the container external network. This is done through iptables rules applied in an init container or the istio CNI. While this should be working inside gVisor information on that is slightly confusing so we need to try it.
Falco and SSH
I am quite sure Falco will not detect ssh traffic in this scenario: it looks for port 22 and ssh* binaries. We have hardly any clusters that make legitimate use of ssh (and they produce 100k+ Falco alerts per day if not silenced). I don't see any legitimate ssh traffic in our agent scenarios which is why we opt for not using it here - this said, of course we would want to use it with Jupyter Notebooks.
Would it make sense to also support setvice account based identities assigned to individual pod? This would help a lot in SPIFFE/Spire setup like in Kagenti to support eg. RFC 8693 based token exchange, and the SVID could be also used in TLS communication with the supervisor (I believe), and no cert-manager would be needed (but Spire then)
and optionally under a gVisor or Kata RuntimeClass that interposes a userspace kernel between the agent and the host.
Can we drop this from the overall framing? There's many possible container isolation tools (including kata, etc) - and for sure enabling use of gVisor should be one big goal, I think a general framing here is more "be able to leverage existing kube isolation tooling" right?
Reacted by Shane UttHi, so I generated my own containerized-agent-wrapper in https://github.com/cgwalters/devaipod (and have reviewed many, many others) but I'd like to invest in OpenShell for some of these things instead.
However to do this, I would like to propose some large scale changes to the podman backend - and I really want this to align with what we do for k8s.
Notes:
- This document has heavy AI generation, but is a result of a fair bit of
interactive design/research work, plus of course big
tip to paude and other projects which already blazed this trail around network proxying - While I know security things quite well (20+ years on that), I'm not a networking person really
- I also didn't really audit all of the communication flows proposed below, so take this as a rough draft
DRAFT: Simplifying OpenShell: Structural Network Isolation via Proxy Sidecar
Motivation
OpenShell's current inner sandboxing — Landlock, seccomp, network namespaces, iptables, and TOFU binary verification, all running inside the sandbox container — is simultaneously too much and too little.
Too restrictive for development agents
Development agents need to install packages (
dnf,apt), experiment with tools, evolve their own environments, and sometimes run nested containers (developing OpenShell in OpenShell). Inner sandboxing breaks all of this. Landlock blocks package manager writes. The custom seccomp profile breaks nested containerization. These are fundamental limitations for the primary use case.Redundant with container runtimes
Because the current "inner" sandboxing needs higher privileges in order to reduce privileges for the inner agent,
it creates duplication with the base configuration already accessible by the container runtime itself!- Landlock overlaps with filesystem isolation configurable in the outer container (
podman run --read-only);
of course, people who want to use Landlock can continue to do so. - Seccomp is already configurable at the container level, and there's well understood tooling for managing that.
- Network namespaces + iptables create a nested namespace inside the container's own namespace just to force traffic through the proxy. This is why the supervisor needs root and
CAP_NET_ADMIN. - TOFU binary verification resolves which binary is making each connection via
/proc/net/tcpand verifies its hash. But if an interpreter is trusted (and agents will use interpreters), it verifies the interpreter binary, not the code being interpreted. The security value is marginal; the complexity cost (CAP_SYS_PTRACE,/procwalking, root) is high.
The fix: structural network isolation, reuse the container runtime
Move the L7 proxy out of the container. Use the container runtime's own network isolation — a Podman
--internalnetwork with no default gateway — as the enforcement boundary. The proxy sidecar sits on both the internal network and the bridge; it's the only route out. The agent can unsetHTTP_PROXYall it wants — there's no route to the internet.Inner sandboxing (Landlock, seccomp) remains available as opt-in defense-in-depth for use cases that need it, but it's no longer required for network isolation. The simpler architecture becomes the default for development agents.
Details
(Click to expand)
Details
Architecture
Each sandbox becomes three Podman resources:
Per sandbox: ┌─ openshell-sbx-{id} network (--internal, dns_enabled=true) ──┐ │ │ │ ┌─────────────────────┐ ┌──────────────────────────────┐ │ │ │ proxy sidecar │ │ agent container │ │ │ │ (openshell-sandbox │ │ (user image) │ │ │ │ OPENSHELL_MODE= │◀─────│ HTTP_PROXY=proxy:3128 │ │ │ │ proxy) │ │ │ │ │ │ - L7 proxy :3128 │ TCP │ openshell-sandbox │ │ │ │ - OPA engine │ fwd │ - SSH/exec relay │ │ │ │ - cred injection │──────│ - gRPC → proxy → gateway │ │ │ │ - inference routing │:8081 │ - opt-in Landlock/seccomp │ │ │ │ - TCP fwd gw:8081 │ │ - no network enforcement │ │ │ │ - singleton MITM CA │ │ - MITM CA via volume (ro) │ │ │ │ - --dns=host resolv │ │ - GIT_SSH_COMMAND disabled │ │ │ └─────────────────────┘ └──────────────────────────────┘ │ │ │ │ └────────┼───────────────────────────────────────────────────────┘ │ also on: openshell bridge network ▼ gateway (host or container, :8081) - sandbox lifecycle via Podman API - gRPC for CLI and supervisor callbacksThe proxy sidecar runs
openshell-sandboxin a newOPENSHELL_MODE=proxy. It handles the L7 proxy, OPA policy, credential injection, and inference routing. It also TCP-forwards the gateway's gRPC port so the agent's supervisor can relay SSH/exec sessions.The agent container runs
openshell-sandboxwithOPENSHELL_PROXY_MODE=sidecar, which tells the supervisor to skip inner network namespace and iptables setup (the topology handles that) while still honoring Landlock/seccomp policy if requested. The agent is connected only to the--internalnetwork — no route to the internet.This mirrors the split-pod model from #981: proxy sidecar = supervisor pod; agent container = agent pod; per-sandbox
--internalnetwork = NetworkPolicy.What this changes
- Network isolation moves to the topology. The
--internalPodman network with no default gateway replaces the inner network namespace + iptables approach. The proxy sidecar is the sole egress path. - Binary identity (TOFU) is removed.
/proc/net/tcpscanning,BinaryIdentityCache,CAP_SYS_PTRACE— fundamentally insecure with LD_PRELOAD and interpreted languages, and can't work across containers. - Proxy moves to sidecar. L7 proxy, OPA engine, credential injection, and inference routing run in a separate container.
- Per-sandbox MITM CAs → singleton. A singleton CA per gateway session is sufficient.
What this preserves
- Inner sandboxing (opt-in). Landlock and seccomp remain available as policy-driven defense-in-depth. They're no longer required for network isolation, but users can opt in for filesystem and syscall restrictions.
- L7 proxy enforcement. OPA policy evaluation, credential injection, inference routing — unchanged in behavior, just moved to the sidecar.
- SSH/exec relay. The agent still runs
openshell-sandboxfor SSH/exec. The gRPC connection routes through the proxy sidecar's TCP port-forward. - Per-sandbox isolation. Each sandbox gets its own internal network. Sandboxes cannot see each other.
- Docker and Kubernetes drivers. Unaffected. This is Podman-only.
Design Decisions
Decision Choice Rationale Isolation model Per-sandbox --internalnetworkSandboxes can't see each other. Mirrors k8s NetworkPolicy from #981. Proxy location Per-sandbox sidecar container Mirrors #981 split-pod. Each sandbox gets its own proxy process. Proxy binary Reuse openshell-sandboxwithOPENSHELL_MODE=proxySupervisor image already has the binary. No new image needed. Agent supervisor OPENSHELL_PROXY_MODE=sidecarSkips inner netns/iptables. Still applies Landlock/seccomp per policy. Still runs SSH/exec. Agent-to-gateway gRPC TCP port-forward through proxy sidecar Agent supervisor connects to proxy's internal IP as if it were the gateway. Binary identity (TOFU) Remove entirely Insecure with LD_PRELOAD; can't work across containers. MITM CA Singleton per gateway session No security benefit to per-sandbox CA. Shared via volume. DNS on internal network dns_enabled: trueContainer-name resolution for debugging. External DNS irrelevant — proxy-aware tools send hostnames via CONNECT. Proxy DNS --dnsset to host resolversDiscovered at driver startup from /run/systemd/resolve/resolv.conf, filtering127.xand169.254.x. Must be set at container creation time.Container IPs Fixed/static per sandbox Proxy= .2, agent=.3. Multi-network requirespodman network connect.Bridge network Existing openshellbridgeProxy sidecars connect here for internet. Dedicated bridge, not default podman.Git-over-SSH Disabled via GIT_SSH_COMMANDPrevents exfiltration. Forces git over HTTPS through proxy. Proxy readiness Sentinel file on shared volume Proxy writes .readyafter CA and listeners are up. Driver polls before starting agent.Podman client Keep hand-rolled client Extend with new methods. Bollard migration is future work.
Prior Art
paude
paude validates the core architecture with a working implementation for Claude Code, Gemini CLI, Cursor, and OpenClaw.
- Per-session
--internalnetwork + proxy sidecar on both networks. No inner sandboxing at all. - Sentinel credentials: agent sees
ANTHROPIC_API_KEY=paude-proxy-managed; proxy swaps for real key in-flight. Equivalent to OpenShell'sSecretResolver. dnsmasqin the proxy sidecar for DNS. Our testing showed this isn't strictly needed — most proxy-aware tools send hostnames via CONNECT.- Fixed IPs (proxy=
.2, agent=.3). Validated. - CA injected via
exec + cat. Our shared-volume approach is cleaner. - Readiness protocol: poll for CA cert before starting agent. Adopted.
alcove
alcove uses a similar dual-network sidecar pattern (Go-based).
- DNS strategy: "proxy resolves." No DNS forwarder — relies on proxy-aware tools sending hostnames in CONNECT. Validated:
curlwithHTTP_PROXYdoes not resolve DNS locally. dns_enabled: trueon internal network for container-name resolution.- SSH disable trick:
GIT_SSH_COMMAND="echo 'SSH disabled' && exit 1". Adopted. - Comprehensive CA trust env vars:
SSL_CERT_FILE,NODE_EXTRA_CA_CERTS,CURL_CA_BUNDLE,GIT_SSL_CAINFO,REQUESTS_CA_BUNDLE. Adopted. - Shared internal network (all sandboxes on one network). We use per-sandbox networks for stronger isolation.
Validated Podman behavior (Podman 5.8.1)
Behavior Result Aardvark-dns resolves external names on --internalNo — NXDOMAIN dns_enabled: trueinternal networkContainer names only, not external Static IPs via --network name:ip=x.x.x.xWorks Multi-network via comma syntax with static IPs Broken — only first network gets IP Multi-network via podman network connectWorks Proxy on both networks Default route via bridge, internet confirmed Agent on internal-only Cannot reach internet (Network unreachable) curlwithHTTP_PROXY, no local DNSWorks — sends CONNECT with hostname wgetwithHTTP_PROXY, no local DNSFails — resolves locally first resolv.confupdated bypodman network connectNo — set at creation time only
Implementation
Phase 0: Remove binary identity (TOFU)
Prep commit. Remove
/proc/net/tcpidentity scanning,BinaryIdentityCache, and all OPA policy fields that depend onbinary_path/binary_sha256. The OPA TCP input simplifies to:{ "host": "api.anthropic.com", "port": 443 }L7 input keeps the request fields but drops
exec:{ "network": { "host": "...", "port": 443 }, "request": { "method": "GET", "path": "/v1/chat", "query_params": {} } }This is a standalone change — it can land independently and unblocks everything else. Key code to remove:
BinaryIdentityCache,find_socket_inode_owners,parse_proc_net_tcp,file_sha256, theentrypoint_pid: Arc<AtomicU32>parameter onProxyHandle::start_with_bind_addr, and theCAP_SYS_PTRACErequirement.Phase 1: Podman client extensions
Add to the hand-rolled Podman client in
crates/openshell-driver-podman/src/client.rs:// Create an --internal bridge network. Idempotent. pub async fn ensure_internal_network(&self, name: &str) -> Result<(), PodmanApiError>; // Inspect container to get its IP on a specific network. pub async fn container_ip(&self, container: &str, network: &str) -> Result<Option<String>, PodmanApiError>; // Connect a running container to an additional network. pub async fn network_connect(&self, network: &str, container: &str) -> Result<(), PodmanApiError>; // Get the /24 subnet base (e.g., "10.89.5") for fixed IP derivation. pub async fn network_subnet_base(&self, name: &str) -> Result<String, PodmanApiError>; // Remove a network. Idempotent. pub async fn remove_network(&self, name: &str) -> Result<(), PodmanApiError>;
Also extend
ContainerSpecwithstatic_ips(per-network) anddns_server(for--dns).Phase 2: Proxy-only mode (
OPENSHELL_MODE=proxy)New mode for
openshell-sandbox— the sidecar entry point. Runs incrates/openshell-sandbox/src/proxy_mode.rs.pub async fn run_proxy_mode(args: SandboxArgs) -> miette::Result<()> { // 1. Connect to gateway, fetch policy, build OPA engine // 2. Fetch provider env, build SecretResolver // 3. Load existing CA from /openshell-tls/ or generate new one // (persist so restarts don't invalidate cached CA in agent) // 4. Build inference context // 5. Start L7 proxy on :3128 // 6. Start TCP port-forward on :8081 → gateway gRPC // 7. Spawn policy poll loop + inference route poller // 8. Write /openshell-tls/.ready (only after BOTH listeners are bound) // 9. Wait for shutdown }
The TCP forwarder is a simple tokio relay. One subtlety:
OPENSHELL_ENDPOINTis an HTTP URL, butTcpStream::connectneedshost:port:async fn run_tcp_forward(listen_addr: SocketAddr, upstream_url: &str) -> miette::Result<()> { let url: url::Url = upstream_url.parse()?; let upstream = format!("{}:{}", url.host_str().unwrap(), url.port_or_known_default().unwrap()); let listener = TcpListener::bind(listen_addr).await?; loop { let (client, _) = listener.accept().await?; let upstream = upstream.clone(); tokio::spawn(async move { if let Ok(server) = TcpStream::connect(&upstream).await { let (mut cr, mut cw) = client.into_split(); let (mut sr, mut sw) = server.into_split(); tokio::select! { _ = tokio::io::copy(&mut cr, &mut sw) => {}, _ = tokio::io::copy(&mut sr, &mut cw) => {}, } } }); } }
Phase 3: Sidecar-aware agent mode (
OPENSHELL_PROXY_MODE=sidecar)When
OPENSHELL_PROXY_MODE=sidecaris set inlib.rs, the supervisor:- Skips inner network namespace, iptables, and proxy startup (the topology and sidecar handle these)
- Still applies Landlock and seccomp if the sandbox policy requests them
- Still starts SSH server and runs the workload
This is distinct from
OPENSHELL_MODE=nestedwhich disables ALL enforcement. Thesidecarmode only disables network-related enforcement.One gotcha:
apply_child_envinssh.rscallsenv_clear()before setting up child processes. The proxy env vars (HTTP_PROXY,SSL_CERT_FILE,GIT_SSH_COMMAND, etc.) set via the container spec will be dropped from SSH/exec sessions. The fix is to propagate a specific set of env vars from the supervisor's own environment into children.Phase 4: Podman driver — per-sandbox sidecar architecture
The sandbox creation sequence in
crates/openshell-driver-podman/src/driver.rs:async fn create_sandbox(&self, sandbox: &DriverSandbox) -> Result<...> { let internal_net = format!("openshell-sbx-{}", sandbox.id); // 1. Per-sandbox internal network self.client.ensure_internal_network(&internal_net).await?; // 2. Fixed IPs from subnet let base = self.client.network_subnet_base(&internal_net).await?; let proxy_ip = format!("{base}.2"); let agent_ip = format!("{base}.3"); // 3. Shared TLS volume (proxy writes CA + .ready; agent reads) let tls_vol = format!("openshell-tls-{}", sandbox.id); self.client.create_volume(&tls_vol).await?; // 4. Proxy sidecar: internal network + fixed IP + host DNS resolvers // (resolv.conf is set at creation time; podman network connect won't update it) let proxy_name = format!("openshell-proxy-{}", sandbox.name); self.client.create_container(&build_proxy_sidecar_spec(...)).await?; // 5. Connect proxy to bridge (gives it internet access via default route) self.client.network_connect(&self.config.network_name, &proxy_name).await?; self.client.start_container(&proxy_name).await?; // 6. Wait for .ready sentinel self.wait_for_proxy_ready(&proxy_name, Duration::from_secs(30)).await?; // 7. Agent: internal-only, fixed IP, OPENSHELL_PROXY_MODE=sidecar self.client.create_container(&build_agent_container_spec(...)).await?; self.client.start_container(&agent_name).await?; Ok(...) }
Deletion is the reverse: agent → proxy → volume → network.
The proxy sidecar container is essentially:
- Image: supervisor image (already has the
openshell-sandboxbinary) - Entrypoint:
/openshell-sandboxwithOPENSHELL_MODE=proxy - Networks: internal only at creation (bridge added via
network_connect) - DNS: explicit
--dnsto host resolvers (discovered from/run/systemd/resolve/resolv.conf, filtering127.xand169.254.x) - Volume: TLS volume mounted rw
The agent container:
- Image: user-specified sandbox image
- Entrypoint: sideloaded
openshell-sandboxvia image volumes - Networks: internal only (no bridge, no internet route)
OPENSHELL_PROXY_MODE=sidecar+ all the proxy/CA env varsGIT_SSH_COMMANDdisabled (prevents git-over-SSH exfiltration)- Volume: TLS volume mounted ro
- Capabilities:
ALLdropped;SYS_ADMINadded only if policy requests Landlock
Agent container env: OPENSHELL_PROXY_MODE=sidecar HTTP_PROXY=http://{proxy_ip}:3128 HTTPS_PROXY=http://{proxy_ip}:3128 OPENSHELL_ENDPOINT=http://{proxy_ip}:8081 SSL_CERT_FILE=/openshell-tls/ca.pem NODE_EXTRA_CA_CERTS=/openshell-tls/ca.pem CURL_CA_BUNDLE=/openshell-tls/ca.pem GIT_SSL_CAINFO=/openshell-tls/ca.pem REQUESTS_CA_BUNDLE=/openshell-tls/ca.pem GIT_SSH_COMMAND=echo 'SSH disabled — use HTTPS' && exit 1Other driver changes:
- Add
openshell.role=proxy/openshell.role=agentlabels. Filterlist_sandboxesonrole=agentto avoid returning duplicates. - Monitor container events for cleanup of orphaned resources.
- Remove
build_container_spec_supervisedandbuild_container_spec_passthrough.
Phase 5: Gateway script and integration tests
Update
gateway-podman.sh. Test the full flow:mise run gateway:podman openshell sandbox create --name test # Verify 3 resources podman network ls | grep openshell-sbx podman ps | grep openshell-proxy-test podman ps | grep openshell-agent-test # Proxy works openshell sandbox exec test -- curl -s https://api.github.com/zen # Bypass fails openshell sandbox exec test -- bash -c 'unset HTTP_PROXY HTTPS_PROXY; curl --connect-timeout 5 https://api.github.com/zen' # → Network unreachable # SSH works openshell sandbox connect test # Git-over-SSH blocked openshell sandbox exec test -- git clone git@github.com:test/repo.git # → SSH disabled — use HTTPS # Cleanup openshell sandbox delete test # No orphaned networks, containers, or volumes
Open Questions
-
wgetand non-proxy-aware tools.wgetresolves DNS locally before connecting to the proxy, so it fails on the internal network. Most development tools (curl, python, node, git) work fine. Known limitation for now; DNS forwarder in the sidecar is future work if needed. -
Proxy code sharing. Both proxy-only mode and inner proxy mode use the same code in
openshell-sandbox. A future refactor could extract shared logic intoopenshell-proxy-core. -
Split OCSF logs. Network/L7 events now come from the proxy sidecar; SSH/process events from the agent supervisor. Two separate log streams. Document as a known change; merging is future work.
-
Proxy readiness: poll vs. healthcheck. Starting with
podman execpolling for.ready. Podman healthcheck is an alternative if polling proves unreliable.
- This document has heavy AI generation, but is a result of a fair bit of
Whether this is implemented or not, I believe a good architectural change that will be beneficial for all possible topologies is to split the
openshell-sandboxcrate from the proxy (Policy enforcement, credential injection) functionality.
I've created an issue for that with a possible implementation strategy: #1305Splitting the sandbox into multiple binaries will unlock whatever topologies will be required in the future, so I believe it is a necessary step in all cases.
I can draft an RFC and/or implement this if it is accepted by the community.Whether this is implemented or not, I believe a good architectural change that will be beneficial for all possible topologies is to split the openshell-sandbox crate from the proxy (Policy enforcement, credential injection) functionality.
I'm not so sure about that. To me, a lot of the current supervisor just overlaps with the base containerization, and if we rely more on that, we don't need it to be standalone at all.
I mean parts of it sure, like landlock stuff or whatever, but there's already several Rust landlock crates.
I'm not so sure about that. To me, a lot of the current supervisor just overlaps with the base containerization, and if we rely more on that, we don't need it to be standalone at all.
I agree, what I'm saying is that there is a benefit in the proxy being a standalone crate & binary, not the supervisor. My reasoning is the same as yours - containerisation already overlaps with what the supervisor is doing, so let's give the freedom to choose your sandboxing technique (plain container, container + gVisor, vm, etc.) but leverage the proxy.
I believe this overlaps with your proposal somewhat. Moving the L7 proxy into a sidecar requires that the proxy is extracted into its own crate & binary.
Hi, gVisor contributor here. Would it be helpful for gVisor to implement Landlock? I ask because I noticed it adds a lot of asterisks to the proposal, so I just want to bring up that one path to resolving that is to just get Landlock done on the gVisor side. There's nothing fundamental stopping it.
I'd point out though that filesystem-level isolation is more powerful when implemented at gVisor's gofer layer (acts as a userspace I/O proxy, think FUSE but with the server half of that living out-of-sandbox). Right now there's just
fsgoferas a plain filesystem access gofer, but there are contributions in-flight to make that more customizable to be able to enforce arbitrary access policies, enable transparent remote storage, shared caching, etc.Reacted by Konstantinos Angelopoulos and Luke Hinds@EtiennePerot , this would be useful - we really need this in nono, as lots of users want to run on gcloud , but instead flip to aws fargate / firecracker due to the lack of landlock support in gvisor - I would even be happy to contribute myself - I created and wrote a lot of the sigstore go code which is widely used by google, so I can vouch it will be the best quality I can get out.
- added a commit that references this issue
on Jun 24, 2026 This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Aug 30, 2026 It looks like #2885 is active work that should probably be marked as closing this issue.
- removedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Aug 31, 2026 Splitting supervisor and agent into separate pods is the right structural move: the control plane should never share fate (or a kernel namespace set) with untrusted workload code. Principle of least privilege for the supervisor pod plus hard isolation for the agent pod, with all cross-talk over explicit APIs. Disclosure: I build vetto, a daemon-less kernel sandbox for AI coding agents (Landlock/seccomp on Linux, Seatbelt on macOS) — mentioning it only because it's directly relevant here. — separation of control and workload is the core of our design too.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Problem Statement
Security concerns as documented in NVIDIA/OpenShell#899.
Proposed Design
RFC: Split Supervisor and Agent into Separate Pods with gVisor Isolation
Status: Draft
Date: 2026-04-26
Authors: @marwinski, @gehoern, @kon-angelo
0. Executive Summary
This proposal extends the Kubernetes compute driver of OpenShell with an alternative deployment mode that significantly strengthens sandbox security while preserving the core OpenShell feature set.
In the current architecture, the supervisor and agent share a single pod. The supervisor runs as root with elevated capabilities — a design that works but creates a wide blast radius if the agent escapes its confinement, and is incompatible with restricted Kubernetes environments such as provided by OpenShift and Gardener.
The proposed split-pod model separates the supervisor and agent into distinct pods. The agent pod runs with zero capabilities, as non-root, and optionally under a gVisor or Kata
RuntimeClassthat interposes a userspace kernel between the agent and the host. A Kubernetes NetworkPolicy restricts the agent's network access to the supervisor's HTTP CONNECT proxy — the same proxy that enforces L7 policy, credential injection, and inference routing today.Key OpenShell features — Landlock filesystem restrictions, SSH access, and the NSSH1 authentication protocol — are retained by injecting a trusted OpenShell binary into the agent pod via the Kubernetes image volumes mechanism. This sideloaded binary runs as the pod entrypoint, configures Landlock (on non-gVisor runtimes), starts the SSH server, and then exec-s the agent process. It requires no elevated capabilities and no cooperation from the agent image. Each feature (sideloading, Landlock, SSH) is independently configurable and can be disabled per sandbox policy.
The existing single-pod (
InPod) mode remains the default and is unchanged. The split-pod model is opt-in, targeting enterprise deployments, managed Kubernetes platforms with restrictive security policies, and environments where stronger workload isolation is required. The only feature lost in the transition is per-binary destination restrictions, which depended on a shared PID namespace between supervisor and agent.1. Motivation and Criticism of Current Architecture
This RFC addresses security concerns with the current single-pod sandbox architecture, as raised in NVIDIA/OpenShell#899 and internal review.
1.1 Current Architecture
Today, the supervisor and agent workload run in a single Kubernetes pod. The supervisor (PID 1) is privileged (
CAP_SYS_ADMIN,CAP_NET_ADMIN,CAP_SYS_PTRACE,CAP_SYSLOG,runAsUser: 0). It creates a network namespace, spawns the agent process inside it, applies Landlock filesystem restrictions, and seccomp filters. All Linux namespaces are shared between supervisor and agent except the network namespace.1.2 Security Concerns
Privilege escalation via shared namespaces. The supervisor runs as root with elevated capabilities in the same pod as the agent. While the agent process itself runs unprivileged, a rogue SUID binary in the container image could allow the agent to escalate to root privileges within the same set of namespaces. The supervisor's elevated capabilities then become reachable.
Implicit trust in the container image. The current model requires trusting the agent container image — at least partially. Landlock restricts filesystem access, but the cluster operator must trust that agent builders do not introduce SUID binaries, world-writable sensitive paths, or other misconfigurations. This is a somewhat shaky trust boundary: the platform claims to sandbox untrusted code, yet relies on the image being well-constructed.
Weak workload isolation. The agent workload runs directly on the host kernel with only seccomp and Landlock as confinement layers. A kernel exploit from the agent process has direct impact on the host. Stronger isolation — such as that provided by gVisor (syscall interception via a userspace kernel) or Kata Containers (lightweight VM boundary) — would significantly reduce the blast radius of a container escape.
Enterprise deployment blockers (Issue #899**).** OpenShift and other managed Kubernetes platforms enforce
restricted-v2SecurityContextConstraints that drop all capabilities and enforcerunAsNonRoot. The current architecture cannot deploy on these platforms without a security exception that most enterprise platform teams are reluctant to grant.Log stream conflation. The supervisor and agent share stdout. OCSF security events, supervisor operational logs, and agent application logs are interleaved in a single stream. A malicious agent could emit lines that mimic OCSF events, poisoning the security audit trail. Splitting into separate pods makes this a non-issue — each pod has its own log stream.
2. Proposed Architecture
2.1 Overview
Split the current single-pod design into two distinct workloads:
Supervisor Pod: Runs the trusted proxy, OPA engine, credential resolver, inference router, and OCSF logger. This is an OpenShell-controlled image.
It is meant to run as the peer to a single agent pod (1:1 sidecar), however we envision that it may also run as a single shared instance (with replicas) serving multiple agent pods.
Agent Pod: Runs the untrusted agent workload. The container image is fully untrusted. The pod runs with no capabilities, as non-root, under a gVisor
RuntimeClass(when available), or sandboxed inside a Kata container. It can only reach the supervisor proxy via Kubernetes NetworkPolicy.2.2 Agent Pod
The agent pod is the untrusted execution environment. Design principles:
restricted-v2SCC on OpenShift.runsc(gVisor's OCI runtime). This interposes a userspace kernel between the agent and the host kernel, providing defense-in-depth against kernel exploits. When gVisor is not available, the pod runs on the standard runtime with reduced isolation (documented trade-off).HTTP_PROXY/HTTPS_PROXYenvironment variables pointing to the supervisor pod's proxy endpoint (e.g.,http://<supervisor-svc>:3128). Cooperative clients honor these variables; non-cooperative traffic is blocked by NetworkPolicy.landlock_restrict_self()is an unprivileged operation. Landlock, SSH, and the sideloaded binary itself are independently configurable and can each be enabled or disabled per sandbox policy.2.3 Supervisor Pod
The supervisor pod is the trusted control plane for one or more agent pods. Initially it appears that a 1:1 relationship is easier as it would remove the necessity for the agent to authenticate, however in environments with many agents a 1:N relationship appears to be preferable. It runs an OpenShell-managed image containing:
openshell:resolve:env:*placeholders in HTTP requests with real secrets fetched from the gateway. Secrets never reach the agent pod.inference.localrequests to configured LLM backends. Enforces inference-specific policy (model allowlists, rate limits).Deployment topologies:
The choice of topology is a deployment configuration, not an architectural constraint. The proxy protocol is the same in both cases.
2.4 Network Policy Enforcement
This NetworkPolicy is the hard enforcement boundary. Even if the agent ignores
HTTP_PROXY, it cannot reach any destination other than the supervisor proxy and cluster DNS. The supervisor then applies L7 policy on top.Important: NetworkPolicy enforcement depends on the CNI plugin. The Kubernetes driver MUST verify that the cluster's CNI supports NetworkPolicy (e.g., Calico, Cilium, Antrea) and warn or fail if it does not.
2.5 Supervisor CA Trust
For TLS interception (MITM), the supervisor generates an ephemeral CA certificate. In the current architecture, this CA is injected into the agent's trust store at process spawn time. In the split-pod model:
REQUESTS_CA_BUNDLEorSSL_CERT_FILEenvironment variables pointing to a mounted cert. No image modification needed.Option C is preferred as it requires no image trust and is purely configuration-driven.
3. What Changes, What Stays, What Is Lost
3.1 Preserved Features
inference.localrouted by proxy to configured backends.no_new_privs). Can be specified as aseccompProfilein the pod'ssecurityContext.3.2 Changed Features
/proc/<pid>/exevia SYS_PTRACEkubectl execavailable as alternative.3.3 Lost Features
Per-binary destination restrictions. In the current model, the proxy resolves which binary opened a network connection by reading
/proc/<pid>/exefrom the supervisor (which shares a PID namespace with the agent). In the split-pod model, the supervisor and agent are in different pods with different PID namespaces. The supervisor proxy sees TCP connections from the agent pod's IP but cannot determine which binary initiated the connection.Possible mitigations:
X-OpenShell-Binary: /usr/bin/curl). The supervisor can use this for advisory policy evaluation but MUST NOT trust it for enforcement — the agent is untrusted and can forge headers.shareProcessNamespace: trueshares PID namespace within a pod, not across pods. This does not help in the split model.Recommendation: Accept the loss of per-binary enforcement. Document it as a known trade-off. The proxy still enforces per-host/port policy and L7 rules for cooperative traffic. NetworkPolicy enforces L3/L4 for all traffic.
4. Landlock in the Split Model
4.1 Landlock Without Root (Unprivileged Containers)
Landlock is explicitly designed for unprivileged use. The
landlock_restrict_self()syscall works underno_new_privsand does not require any Linux capability. However, the current OpenShell implementation uses a two-phase approach:PathFdhandles to all allowed paths. Currently runs as root to access paths the agent user cannot read (e.g.,/usr/sbin).restrict_self(). Does not require root.In an unprivileged container, Phase 1 can only open paths readable by the container's UID. Paths outside the UID's access will fail. The existing
best_effortmode handles this gracefully — inaccessible paths are skipped with a warning, and Landlock is applied for the paths that could be opened.Verdict: Landlock works in unprivileged containers with degraded coverage (best-effort mode). With the image volume sideloading approach (§2.2), the platform injects the OpenShell binary into the agent pod as the entrypoint. This binary sets up Landlock before exec-ing the agent process — the same flow as the current single-pod model, minus root. No action required from the agent deployer.
4.2 Landlock Under gVisor
gVisor implements its own syscall table in a userspace kernel (the "sentry"). As of the current gVisor release, gVisor does not implement the Landlock syscalls (
landlock_create_ruleset,landlock_add_rule,landlock_restrict_self). These syscalls will returnENOSYSinside a gVisor-sandboxed container.This means Landlock cannot be used inside gVisor pods. However, this is not necessarily a problem:
runscOCI runtime supports restricting filesystem access via OCI spec mounts (read-only, masked paths, etc.).4.3 Recommendation
Landlock is an optional, platform-managed feature. On non-gVisor runtimes, the platform injects the OpenShell sideloaded binary via Kubernetes image volumes. This binary applies Landlock rules at startup (best-effort mode) before exec-ing the agent process. The agent deployer does not need to include any OpenShell components in their image — the platform handles injection transparently. The deployer can disable Landlock via sandbox policy if it interferes with their workload.
For gVisor pods, Landlock is unavailable (gVisor does not implement the Landlock syscalls). Filesystem isolation is configured via the pod spec and gVisor's runtime configuration instead. This is controlled by the platform, not the agent image.
5. SSH Access
5.1 Current Mechanism
The supervisor embeds an SSH server (russh) that listens on port 2222. Users authenticate via the NSSH1 handshake protocol (time-bounded shared secret). The SSH session spawns a shell in the agent's network namespace with dropped privileges.
5.2 Sideloaded SSH in Split Model
With the image volume sideloading approach (§2.2), the SSH server is part of the sideloaded OpenShell binary injected into the agent pod by the platform. The binary runs as the pod entrypoint, starts the SSH server, and then exec-s the agent process. This preserves the current NSSH1 authentication flow and user experience without requiring the agent deployer to include any OpenShell components in their image.
Key properties:
5.3 Alternatives
kubectl exec(fallback). When SSH is disabled, users access the agent viakubectl execor the Kubernetes API. This is the standard Kubernetes-native approach — no image modification, no additional network exposure. However, it requires Kubernetes RBAC and lacks the NSSH1 authentication protocol.5.4 Recommendation
Enable SSH via the sideloaded binary by default. Deployers can opt out via sandbox policy (
ssh: disabled). When SSH is disabled,kubectl execis the fallback. Document that SSH in the cluster may trigger Falco alerts and provide guidance on configuring exceptions.6. Observability Benefits
Splitting supervisor and agent into separate pods provides clear observability wins:
6.1 Log Stream Separation
A log shipper (Fluentd, Fluent Bit, Vector) can collect each stream independently. OCSF events cannot be spoofed by the agent — they originate from a different pod with a different log source identifier. This eliminates the log poisoning attack vector where a malicious agent emits fake OCSF events to its stdout.
6.2 gVisor + Falco Integration
gVisor exposes a runtime monitoring interface that Falco can consume. With the agent running under gVisor, Falco can inspect:
This provides the same class of runtime security monitoring that Falco provides for standard containers, but with the added isolation boundary of gVisor's userspace kernel.
6.3 Resource Attribution
Separate pods enable accurate resource metering per agent workload:
7. Agent Identity in Shared Supervisor Topology
In the 1:N topology (one supervisor serving multiple agents), the supervisor must identify which agent is making each proxy request.
7.1 Identification Mechanisms
Source IP mapping. The supervisor maintains a mapping of agent pod IPs to agent identities. When a CONNECT request arrives from
10.0.1.42, the supervisor looks up which agent pod owns that IP. This mapping is maintained via the Kubernetes API (watch pod events) or via the gateway's sandbox registry.Client certificate (mTLS). Each agent pod is provisioned with a unique client certificate (via Kubernetes Secrets or cert-manager). The agent's
HTTP_PROXYconfiguration includes the client cert. The supervisor validates the cert and extracts the agent identity from the CN/SAN.Token header. A per-agent bearer token is injected as an environment variable. The agent's HTTP client includes it in a
Proxy-Authorizationheader. Simpler than mTLS but less robust (token can be exfiltrated by the agent and replayed from a different context).7.2 Recommendation
Use token header as the primary identification mechanism as it is part of the standard.
8. Deployment and Migration
8.1 Backward Compatibility
This RFC does NOT remove the existing single-pod (
InPod) deployment mode. The current architecture remains the default for:openshell bootstrapThe split-pod model is a new
NetworkModevariant (e.g.,PlatformorSplitPod) selected via configuration.8.2 Configuration
The
sideloadsetting controls whether the platform injects the OpenShell binary into the agent pod via Kubernetes image volumes. When enabled, the binary acts as the pod entrypoint and can manage Landlock and SSH. When disabled, the agent pod runs the deployer's image entrypoint directly — Landlock and SSH are unavailable, and the pod relies solely on NetworkPolicy, gVisor, and pod spec restrictions for isolation.sshandlandlockare independently togglable but both requiresideload: enabled. Settingsideload: disabledimplicitly disables both.8.3 Kubernetes Driver Changes
The Kubernetes driver needs to:
network_mode: platform.HTTP_PROXY,HTTPS_PROXY,NO_PROXY,SSL_CERT_FILE).sideload: enabled. Override the agent container's command to the sideloaded binary path. Configure the binary with the appropriate Landlock and SSH settings.9. Security Analysis
9.1 Threat Model Comparison
/proc. Spoofable via symlinks but detectable.9.2 Trust Boundaries
9.3 Key Security Property
No mistake by the agent deployer can compromise the overall system. The agent image is untrusted. The agent pod has no capabilities. Network access is restricted by platform-enforced NetworkPolicy. Secrets are resolved in the supervisor, never exposed to the agent. gVisor provides kernel-level isolation. The only cooperation expected from the agent is honoring
HTTP_PROXY— and non-cooperation is handled by NetworkPolicy denial, not by trusting the agent to behave.10. Open Questions
runscruntime to be installed on nodes and aRuntimeClassto be defined. How do we handle clusters where gVisor is not available? Fall back to standard runtime with documented reduced isolation?11. Summary
This RFC proposes splitting the OpenShell sandbox into a trusted supervisor pod and an untrusted agent pod, connected via Kubernetes NetworkPolicy and an HTTP CONNECT proxy. The agent pod can optionally run under gVisor for strong kernel-level isolation.
The key trade-offs:
The existing
InPodmode is preserved as the default for environments where gVisor is unavailable and elevated capabilities are acceptable. The split-pod model is an opt-in deployment mode for enterprise and security-sensitive environments.Alternatives Considered
We did compare it with the current design.
Agent Investigation
We used an agent to investigate the codebase and prepare the design. We also did a lot of"human" review that makes us believe that the proposal is plausible and implementable.
We are ready and willing to implement this proposal.
Checklist