Repository navigation
Kubernetes driver never requests CAP_SETPCAP, so combined-topology sandboxes crash-loop with EPERM on privilege drop #2751
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Aug 14, 2026 Hello, I tried to reproduce on Kubernetes, but I did not manage to hit the same issue.
Given that you use
ocin your steps, could be this is not general Kubernetes issue, but it affects OpenShift specifically?I am on macOS, so I used a
lima-managed VM, like this❯ limactl start --name k8s --cpus 4 --memory 8 --disk 40 template:k8s ? Creating an instance `k8s` Proceed with the current configuration INFO[0011] Starting the instance `k8s` with internal VM driver `vz` INFO[0011] Attempting to download the image arch=aarch64 digest="sha256:7bcf159e29ad0000bfed9c57875908c39268f5ed1257f4958fa6a9f5f60edd54" location="https://cloud-images.ubuntu.com/releases/resolute/release-20260720/ubuntu-26.04-server-cloudimg-arm64.img" Downloading the image (ubuntu-26.04-server-cloudimg-arm64.img) ... INFO[0891] READY. Run `limactl shell k8s` to open the shell. INFO[0891] Message from the instance `k8s`: To run `kubectl` on the host (assumes kubectl is installed), run the following commands: ------ export KUBECONFIG="/Users/jdanek/.lima/k8s/copied-from-guest/kubeconfig.yaml" kubectl ... ------Then I did the necessary setup, I used rancher local-path-provisioner to get a necessary PVC:
❯ kubectl apply -f https://github.com/kubernetes-sigs/agent-sandbox/releases/download/v0.5.6/sandbox.yaml ❯ kubectl apply -f https://raw.githubusercontent.com/rancher/local-path-provisioner/v0.0.33/deploy/local-path-storage.yaml ❯ kubectl patch storageclass local-path -p '{"metadata": {"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'(the
latesturl for agent-standbox did not work, so I put found the latest version string and used that)And I installed openshell with the
combinedtopology❯ CHART=0.0.110 ❯ helm upgrade --install openshell \ oci://ghcr.io/nvidia/openshell/helm-chart \ --version "$CHART" \ --namespace openshell \ --set supervisor.topology=combined \ --set server.disableTls=true \ --set server.auth.allowUnauthenticatedUsers=true Release "openshell" does not exist. Installing it now. Pulled: ghcr.io/nvidia/openshell/helm-chart:0.0.110 Digest: sha256:8bf947365a46184e979a30039a7f22fe57e150827ab699bc688e9260f7c4af12 NAME: openshell LAST DEPLOYED: Fri Aug 21 17:58:05 2026 NAMESPACE: openshell STATUS: deployed REVISION: 1 DESCRIPTION: Install complete TEST SUITE: NoneWhich seems to have been applied
❯ kubectl -n openshell get cm -o yaml | grep -n -A2 -E 'topology' 70: topology = "combined"Then, trying to run a pod, the pod started just fine for me, baring some hiccup with slow network downloading sandbox pod; the cli timeouted, but the sandbox did come up
❯ kubectl -n openshell port-forward svc/openshell 8080:8080 Forwarding from 127.0.0.1:8080 -> 8080 Forwarding from [::1]:8080 -> 8080 Handling connection for 8080 Handling connection for 8080 Handling connection for 8080❯ openshell gateway add http://127.0.0.1:8080 --local --name k8s ✓ Gateway 'k8s' added and set as active Endpoint: http://127.0.0.1:8080 Type: local Auth: plaintext ❯ openshell sandbox create --name cap-setpcap-repro -- true Created sandbox: cap-setpcap-repro ✗ sandbox provisioning timed out after 300s. Last reported status: DependenciesNotReady: Pod exists with phase: Pending ✓ Sandbox allocated (4s) ✓ Image pulled (13 MB) (6s) Error: × sandbox provisioning timed out after 300s. Last reported status: DependenciesNotReady: Pod exists with phase: Pending ❯ openshell sandbox create --name cap-setpcap-repro -- true Error: × sandbox 'cap-setpcap-repro' already exists │ │ hint: delete it first with: openshell sandbox delete <name> │ or use a different name ❯ kubectl -n openshell get sandboxes NAME READY REASON AGE default--cap-setpcap-repro True DependenciesReady 12m❯ kubectl get pod -n openshell -o yaml default--cap-setpcap-repro
apiVersion: v1 kind: Pod metadata: annotations: agents.x-k8s.io/propagated-annotations: openshell.io/sandbox-id openshell.io/sandbox-id: fb863998-78c4-49e8-bbee-b944471864e2 creationTimestamp: "2026-08-21T16:04:15Z" generation: 1 labels: agents.x-k8s.io/sandbox-name-hash: 1fd73180 name: default--cap-setpcap-repro namespace: openshell ownerReferences: - apiVersion: agents.x-k8s.io/v1beta1 blockOwnerDeletion: true controller: true kind: Sandbox name: default--cap-setpcap-repro uid: 820af084-1316-41c8-87da-aa984ca61923 resourceVersion: "3942" uid: c402b511-f1e5-4c08-8337-2f35b76d3683 spec: automountServiceAccountToken: false containers: - command: - /opt/openshell/bin/openshell-sandbox - --workdir - /sandbox env: - name: OPENSHELL_SANDBOX_ID value: fb863998-78c4-49e8-bbee-b944471864e2 - name: OPENSHELL_SANDBOX value: cap-setpcap-repro - name: OPENSHELL_ENDPOINT value: http://openshell.openshell.svc.cluster.local:8080 - name: OPENSHELL_SANDBOX_COMMAND value: sleep infinity - name: OPENSHELL_TELEMETRY_ENABLED value: "true" - name: OPENSHELL_SSH_SOCKET_PATH value: /run/openshell/ssh.sock - name: OPENSHELL_K8S_SA_TOKEN_FILE value: /var/run/secrets/openshell/token - name: OPENSHELL_OCI_IMAGE_USER - name: OPENSHELL_SANDBOX_UID value: "1000" - name: OPENSHELL_SANDBOX_GID value: "1000" image: ghcr.io/nvidia/openshell-community/sandboxes/base:latest imagePullPolicy: Always name: agent resources: {} securityContext: appArmorProfile: type: Unconfined capabilities: add: - SYS_ADMIN - NET_ADMIN - SYS_PTRACE - SYSLOG runAsUser: 0 terminationMessagePath: /dev/termination-log terminationMessagePolicy: File volumeMounts: - mountPath: /var/run/secrets/openshell name: openshell-sa-token readOnly: true - mountPath: /opt/openshell/bin name: openshell-supervisor-bin readOnly: true - mountPath: /sandbox name: workspace dnsPolicy: ClusterFirst enableServiceLinks: true initContainers: - command: - sh - -c - if [ ! -f /workspace-pvc/.workspace-initialized ]; then if [ -d /sandbox ]; then tmp=$(mktemp) && rm -f "$tmp" && (cd /sandbox && find . -mindepth 1 -maxdepth 1 -exec tar -cf "$tmp" {} +) && if [ -f "$tmp" ]; then tar -C /workspace-pvc --no-same-owner --no-same-permissions --touch -xf "$tmp" && rm -f "$tmp"; fi; fi && touch /workspace-pvc/.workspace-initialized; fi image: ghcr.io/nvidia/openshell-community/sandboxes/base:latest imagePullPolicy: Always name: workspace-init resources: {} securityContext: runAsUser: 0 terminationMessagePath: /dev/termination-log terminationMessagePolicy: File volumeMounts: - mountPath: /workspace-pvc name: workspace nodeName: lima-k8s preemptionPolicy: PreemptLowerPriority priority: 0 restartPolicy: Always schedulerName: default-scheduler securityContext: fsGroup: 1000 serviceAccount: openshell-sandbox serviceAccountName: openshell-sandbox terminationGracePeriodSeconds: 30 tolerations: - effect: NoExecute key: node.kubernetes.io/not-ready operator: Exists tolerationSeconds: 300 - effect: NoExecute key: node.kubernetes.io/unreachable operator: Exists tolerationSeconds: 300 volumes: - name: openshell-sa-token projected: defaultMode: 256 sources: - serviceAccountToken: audience: openshell-gateway expirationSeconds: 3600 path: token - image: pullPolicy: IfNotPresent reference: ghcr.io/nvidia/openshell/supervisor:7909fb5d0f54a06e26eb79e47885d7dd105aef24 name: openshell-supervisor-bin - name: workspace persistentVolumeClaim: claimName: workspace-default--cap-setpcap-repro status: conditions: - lastProbeTime: null lastTransitionTime: "2026-08-21T16:04:19Z" observedGeneration: 1 status: "True" type: PodReadyToStartContainers - lastProbeTime: null lastTransitionTime: "2026-08-21T16:13:14Z" observedGeneration: 1 status: "True" type: Initialized - lastProbeTime: null lastTransitionTime: "2026-08-21T16:13:15Z" observedGeneration: 1 status: "True" type: Ready - lastProbeTime: null lastTransitionTime: "2026-08-21T16:13:15Z" observedGeneration: 1 status: "True" type: ContainersReady - lastProbeTime: null lastTransitionTime: "2026-08-21T16:04:19Z" observedGeneration: 1 status: "True" type: PodScheduled containerStatuses: - containerID: containerd://cf81decc4cd667dfdf3515868c75e869899cb0a22386e3335982a9af4ad2b6f0 image: ghcr.io/nvidia/openshell-community/sandboxes/base:latest imageID: ghcr.io/nvidia/openshell-community/sandboxes/base@sha256:aeef1c63f00e2913ea002ccb3aaf925f338b5c5d70e63576f0d95c16a138044e lastState: {} name: agent ready: true resources: {} restartCount: 0 started: true state: running: startedAt: "2026-08-21T16:13:15Z" user: linux: gid: 0 supplementalGroups: - 0 - 1000 uid: 0 volumeMounts: - mountPath: /var/run/secrets/openshell name: openshell-sa-token readOnly: true recursiveReadOnly: Disabled - mountPath: /opt/openshell/bin name: openshell-supervisor-bin readOnly: true recursiveReadOnly: Disabled - mountPath: /sandbox name: workspace hostIP: 192.168.5.15 hostIPs: - ip: 192.168.5.15 initContainerStatuses: - containerID: containerd://00438f907a9c9c556fe6e921b3f41554f7f4e9d47c0a045e9543c3b27dc08d9e image: ghcr.io/nvidia/openshell-community/sandboxes/base:latest imageID: ghcr.io/nvidia/openshell-community/sandboxes/base@sha256:aeef1c63f00e2913ea002ccb3aaf925f338b5c5d70e63576f0d95c16a138044e lastState: {} name: workspace-init ready: true resources: {} restartCount: 0 started: false state: terminated: containerID: containerd://00438f907a9c9c556fe6e921b3f41554f7f4e9d47c0a045e9543c3b27dc08d9e exitCode: 0 finishedAt: "2026-08-21T16:13:13Z" reason: Completed startedAt: "2026-08-21T16:13:13Z" user: linux: gid: 0 supplementalGroups: - 0 - 1000 uid: 0 volumeMounts: - mountPath: /workspace-pvc name: workspace observedGeneration: 1 phase: Running podIP: 10.244.0.10 podIPs: - ip: 10.244.0.10 qosClass: BestEffort resources: {} startTime: "2026-08-21T16:04:19Z"The above includes
securityContext: appArmorProfile: type: Unconfined capabilities: add: - SYS_ADMIN - NET_ADMIN - SYS_PTRACE - SYSLOG runAsUser: 0For the openshell-0 pod
❯ kubectl get pod -n openshell -o yaml openshell-0
apiVersion: v1 kind: Pod metadata: annotations: checksum/gateway-config: 8d5fbcf2a0e531b1f7b2b998f6879a261b6284d0bef6c1bf2c44ee5c6d3897f8 creationTimestamp: "2026-08-21T15:58:29Z" generateName: openshell- generation: 1 labels: app.kubernetes.io/instance: openshell app.kubernetes.io/managed-by: Helm app.kubernetes.io/name: openshell app.kubernetes.io/version: 0.0.110 apps.kubernetes.io/pod-index: "0" controller-revision-hash: openshell-5974c7bcbd helm.sh/chart: helm-chart-0.0.110 statefulset.kubernetes.io/pod-name: openshell-0 name: openshell-0 namespace: openshell ownerReferences: - apiVersion: apps/v1 blockOwnerDeletion: true controller: true kind: StatefulSet name: openshell uid: 76429d46-5f79-4592-a7bc-43832691d866 resourceVersion: "2765" uid: 2c4953ec-585f-4934-9283-af1859516b07 spec: containers: - args: - --config - /etc/openshell/gateway.toml - --db-url - sqlite:/var/openshell/openshell.db env: - name: OPENSHELL_GATEWAY_CREDENTIAL_KEY_ENCRYPTION_KEY valueFrom: secretKeyRef: key: key-encryption-key name: openshell-credential-storage-key-encryption-key - name: OPENSHELL_TELEMETRY_ENABLED value: "true" image: ghcr.io/nvidia/openshell/gateway:0.0.110 imagePullPolicy: IfNotPresent livenessProbe: failureThreshold: 3 httpGet: path: /healthz port: health scheme: HTTP initialDelaySeconds: 2 periodSeconds: 5 successThreshold: 1 timeoutSeconds: 1 name: openshell-gateway ports: - containerPort: 8080 name: grpc protocol: TCP - containerPort: 8081 name: health protocol: TCP - containerPort: 9090 name: metrics protocol: TCP readinessProbe: failureThreshold: 3 httpGet: path: /readyz port: health scheme: HTTP initialDelaySeconds: 1 periodSeconds: 2 successThreshold: 1 timeoutSeconds: 1 resources: {} securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsNonRoot: true runAsUser: 1000 startupProbe: failureThreshold: 30 httpGet: path: /healthz port: health scheme: HTTP periodSeconds: 2 successThreshold: 1 timeoutSeconds: 1 terminationMessagePath: /dev/termination-log terminationMessagePolicy: File volumeMounts: - mountPath: /var/openshell name: openshell-data - mountPath: /etc/openshell name: gateway-config readOnly: true - mountPath: /etc/openshell-jwt name: sandbox-jwt readOnly: true - mountPath: /var/run/secrets/kubernetes.io/serviceaccount name: kube-api-access-zrv4r readOnly: true dnsPolicy: ClusterFirst enableServiceLinks: true hostname: openshell-0 nodeName: lima-k8s preemptionPolicy: PreemptLowerPriority priority: 0 restartPolicy: Always schedulerName: default-scheduler securityContext: fsGroup: 1000 serviceAccount: openshell serviceAccountName: openshell subdomain: openshell terminationGracePeriodSeconds: 5 tolerations: - effect: NoExecute key: node.kubernetes.io/not-ready operator: Exists tolerationSeconds: 300 - effect: NoExecute key: node.kubernetes.io/unreachable operator: Exists tolerationSeconds: 300 volumes: - name: openshell-data persistentVolumeClaim: claimName: openshell-data-openshell-0 - configMap: defaultMode: 420 name: openshell-config name: gateway-config - name: sandbox-jwt secret: defaultMode: 256 secretName: openshell-jwt-keys - name: kube-api-access-zrv4r projected: defaultMode: 420 sources: - serviceAccountToken: expirationSeconds: 3607 path: token - configMap: items: - key: ca.crt path: ca.crt name: kube-root-ca.crt - downwardAPI: items: - fieldRef: apiVersion: v1 fieldPath: metadata.namespace path: namespace status: conditions: - lastProbeTime: null lastTransitionTime: "2026-08-21T16:03:19Z" observedGeneration: 1 status: "True" type: PodReadyToStartContainers - lastProbeTime: null lastTransitionTime: "2026-08-21T16:03:18Z" observedGeneration: 1 status: "True" type: Initialized - lastProbeTime: null lastTransitionTime: "2026-08-21T16:03:21Z" observedGeneration: 1 status: "True" type: Ready - lastProbeTime: null lastTransitionTime: "2026-08-21T16:03:21Z" observedGeneration: 1 status: "True" type: ContainersReady - lastProbeTime: null lastTransitionTime: "2026-08-21T16:03:18Z" observedGeneration: 1 status: "True" type: PodScheduled containerStatuses: - containerID: containerd://50aa7e3ca3af8162adc186b7a6d078b63216c2ae6fc32aa60af2f232e7e1aa4e image: ghcr.io/nvidia/openshell/gateway:0.0.110 imageID: ghcr.io/nvidia/openshell/gateway@sha256:398bf373bc8cf23bd9b1b1096bf5ecda814ff459c2f335ecc9c27aacdcef9fa1 lastState: {} name: openshell-gateway ready: true resources: {} restartCount: 0 started: true state: running: startedAt: "2026-08-21T16:03:19Z" user: linux: gid: 0 supplementalGroups: - 0 - 1000 uid: 1000 volumeMounts: - mountPath: /var/openshell name: openshell-data - mountPath: /etc/openshell name: gateway-config readOnly: true recursiveReadOnly: Disabled - mountPath: /etc/openshell-jwt name: sandbox-jwt readOnly: true recursiveReadOnly: Disabled - mountPath: /var/run/secrets/kubernetes.io/serviceaccount name: kube-api-access-zrv4r readOnly: true recursiveReadOnly: Disabled hostIP: 192.168.5.15 hostIPs: - ip: 192.168.5.15 observedGeneration: 1 phase: Running podIP: 10.244.0.8 podIPs: - ip: 10.244.0.8 qosClass: BestEffort resources: {} startTime: "2026-08-21T16:03:18Z"That includes
securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsNonRoot: true runAsUser: 1000I see no crashloops or any errors besides seemingly innocuous
❯ kubectl logs -n openshell default--cap-setpcap-repro -c agent 2026-08-21T16:13:16.071Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/usr/local/bin/opencode container_path=/proc/38/root/usr/local/bin/opencode pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE. 2026-08-21T16:13:16.072Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/usr/local/bin/opencode container_path=/proc/38/root/usr/local/bin/opencode pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE. 2026-08-21T16:13:16.072Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/usr/bin/wget container_path=/proc/38/root/usr/bin/wget pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE. 2026-08-21T16:13:16.079Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/usr/bin/wget container_path=/proc/38/root/usr/bin/wget pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE. 2026-08-21T16:13:16.079Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/app/.venv/bin/python container_path=/proc/38/root/app/.venv/bin/python pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE. 2026-08-21T16:13:16.079Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/app/.venv/bin/python3 container_path=/proc/38/root/app/.venv/bin/python3 pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE. 2026-08-21T16:13:16.079Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/app/.venv/bin/pip container_path=/proc/38/root/app/.venv/bin/pip pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE.📋 triage-agent
Closing as obsolete after RFC 0012. The failure is specific to the retired Kubernetes combined topology, where the privileged supervisor ran in the workload container. Kubernetes now uses separate supervisor and workload pods through #2942/#3144, so adding
CAP_SETPCAPto that removed topology would restore privilege to a design OpenShell no longer uses.- removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Oct 1, 2026
Summary
On Kubernetes,
topology = "combined"sandboxes always crash on startup withEPERMbecause the supervisor's privilege-drop routine requiresCAP_SETPCAP, which the Kubernetes driver never requests in the pod'ssecurityContext.capabilities.addlist — under any config.Root cause
drop_privileges_with_identity(crates/openshell-supervisor-process/src/process.rs) callsdrop_capability_bounding_set()beforesetuid()/setgid(), wheneverenforcement_mode.uses_privileged_process_setup()is true (i.e.topology = "combined"→ProcessEnforcementMode::Full, seecrates/openshell-sandbox/src/lib.rs:857-858).drop_capability_bounding_set()(process.rs:277-286) callscapctl::caps::bounding::clear(), which requiresCAP_SETPCAP.crates/openshell-driver-kubernetes/src/driver.rs:2692-2699) only ever addsSYS_ADMIN, NET_ADMIN, SYS_PTRACE, SYSLOG(plusSETUID, SETGID, DAC_READ_SEARCHwhenenable_user_namespacesis on) —SETPCAPis never added, in any code path.capabilities: { drop: ["ALL"] }unconditionally, the supervisor (running as root,runAsUser: 0) never actually holdsCAP_SETPCAP, so the bounding-set clear fails withEPERMand the pod crash-loops before ever exec'ing the workload.By contrast, the Podman driver gets this right:
crates/openshell-driver-podman/container.rs:1048-1051explicitly addsSETPCAPtocap_addwith a comment referencing exactly this requirement.Reproduction
topology = "combined"(the documented/default topology) on any cluster where sandbox pods run under a restricted PodSecurityStandard / custom SCC that only grants the capabilities the driver actually requests (i.e. not a blanketprivilegedSCC).CreateSandbox.agentcontainer crash-loop with:oc logs <pod> -c agentshows nothing more specific — the error surfaces fromvalidate_capability_bounding_set_clear(process.rs:289-311) with a non-empty remaining bounding set, which does not hit the "already empty" fallback path.We also confirmed the network-policy portion of sandbox setup (which needs
SYS_ADMIN/NET_ADMIN) succeeds and reportsReportPolicyStatus: status=loadedbefore this crash — this is specifically the privilege-drop step, not the network-policy setup step.Notes
docs/kubernetes/openshift.mdx:12) note the OpenShift install path is "experimental" and recommend theprivilegedSCC — but since the driver unconditionally setscapabilities: { drop: ["ALL"], add: [...] }regardless of SCC, granting a more permissive SCC doesn't help; the pod's own requested capability list is still missingSETPCAP.enable_user_namespaces = truedoes NOT fix this — it only addsSETUID/SETGID/DAC_READ_SEARCH, notSETPCAP.topology = "sidecar"(ProcessEnforcementMode::NetworkOnly), which skipsdrop_privilegesentirely since the container already runs as the target UID.Suggested fix
Add
"SETPCAP"to the capability list built incrates/openshell-driver-kubernetes/src/driver.rs(around line 2692), mirroring the Podman driver'scap_addlist, for any topology whoseProcessEnforcementMode::uses_privileged_process_setup()is true.Environment
devbranch, commitc4b500a(2026-08-14)SYS_ADMIN, NET_ADMIN, NET_RAW, SYS_PTRACE, SYSLOG, CHOWN, FOWNER, DAC_READ_SEARCH, SETUID, SETGID(noSETPCAP)