Skip to content

Kubernetes driver never requests CAP_SETPCAP, so combined-topology sandboxes crash-loop with EPERM on privilege drop #2751

Description

@thoraxe

Summary

On Kubernetes, topology = "combined" sandboxes always crash on startup with EPERM because the supervisor's privilege-drop routine requires CAP_SETPCAP, which the Kubernetes driver never requests in the pod's securityContext.capabilities.add list — under any config.

Root cause

  • drop_privileges_with_identity (crates/openshell-supervisor-process/src/process.rs) calls drop_capability_bounding_set() before setuid()/setgid(), whenever enforcement_mode.uses_privileged_process_setup() is true (i.e. topology = "combined" → ProcessEnforcementMode::Full, see crates/openshell-sandbox/src/lib.rs:857-858).
  • drop_capability_bounding_set() (process.rs:277-286) calls capctl::caps::bounding::clear(), which requires CAP_SETPCAP.
  • The Kubernetes driver's capability list (crates/openshell-driver-kubernetes/src/driver.rs:2692-2699) only ever adds SYS_ADMIN, NET_ADMIN, SYS_PTRACE, SYSLOG (plus SETUID, SETGID, DAC_READ_SEARCH when enable_user_namespaces is on) — SETPCAP is never added, in any code path.
  • Since the driver also sets capabilities: { drop: ["ALL"] } unconditionally, the supervisor (running as root, runAsUser: 0) never actually holds CAP_SETPCAP, so the bounding-set clear fails with EPERM and the pod crash-loops before ever exec'ing the workload.

By contrast, the Podman driver gets this right: crates/openshell-driver-podman/container.rs:1048-1051 explicitly adds SETPCAP to cap_add with a comment referencing exactly this requirement.

Reproduction

  1. Deploy the Kubernetes driver with topology = "combined" (the documented/default topology) on any cluster where sandbox pods run under a restricted PodSecurityStandard / custom SCC that only grants the capabilities the driver actually requests (i.e. not a blanket privileged SCC).
  2. Create a sandbox via CreateSandbox.
  3. Observe the agent container crash-loop with:
    Error:   × EPERM: Operation not permitted
    
  4. oc logs <pod> -c agent shows nothing more specific — the error surfaces from validate_capability_bounding_set_clear (process.rs:289-311) with a non-empty remaining bounding set, which does not hit the "already empty" fallback path.

We also confirmed the network-policy portion of sandbox setup (which needs SYS_ADMIN/NET_ADMIN) succeeds and reports ReportPolicyStatus: status=loaded before this crash — this is specifically the privilege-drop step, not the network-policy setup step.

Notes

  • Docs (docs/kubernetes/openshift.mdx:12) note the OpenShift install path is "experimental" and recommend the privileged SCC — but since the driver unconditionally sets capabilities: { drop: ["ALL"], add: [...] } regardless of SCC, granting a more permissive SCC doesn't help; the pod's own requested capability list is still missing SETPCAP.
  • enable_user_namespaces = true does NOT fix this — it only adds SETUID/SETGID/DAC_READ_SEARCH, not SETPCAP.
  • Workaround we're using in the meantime: switching to topology = "sidecar" (ProcessEnforcementMode::NetworkOnly), which skips drop_privileges entirely since the container already runs as the target UID.

Suggested fix

Add "SETPCAP" to the capability list built in crates/openshell-driver-kubernetes/src/driver.rs (around line 2692), mirroring the Podman driver's cap_add list, for any topology whose ProcessEnforcementMode::uses_privileged_process_setup() is true.

Environment

  • OpenShell built from dev branch, commit c4b500a (2026-08-14)
  • OpenShift, custom SCC granting SYS_ADMIN, NET_ADMIN, NET_RAW, SYS_PTRACE, SYSLOG, CHOWN, FOWNER, DAC_READ_SEARCH, SETUID, SETGID (no SETPCAP)

Activity

  1. jiridanek commented on Aug 21, 2026

    @jiridanek

    Hello, I tried to reproduce on Kubernetes, but I did not manage to hit the same issue.

    Given that you use oc in your steps, could be this is not general Kubernetes issue, but it affects OpenShift specifically?

    I am on macOS, so I used a lima-managed VM, like this

    ❯ limactl start --name k8s --cpus 4 --memory 8 --disk 40 template:k8s  
    
    ? Creating an instance `k8s` Proceed with the current configuration
    INFO[0011] Starting the instance `k8s` with internal VM driver `vz` 
    INFO[0011] Attempting to download the image              arch=aarch64 digest="sha256:7bcf159e29ad0000bfed9c57875908c39268f5ed1257f4958fa6a9f5f60edd54" location="https://cloud-images.ubuntu.com/releases/resolute/release-20260720/ubuntu-26.04-server-cloudimg-arm64.img"
    Downloading the image (ubuntu-26.04-server-cloudimg-arm64.img)
    ...
    INFO[0891] READY. Run `limactl shell k8s` to open the shell. 
    INFO[0891] Message from the instance `k8s`:             
    
    To run `kubectl` on the host (assumes kubectl is installed), run the following commands:
    ------
    export KUBECONFIG="/Users/jdanek/.lima/k8s/copied-from-guest/kubeconfig.yaml"
    kubectl ...
    ------
    
    

    Then I did the necessary setup, I used rancher local-path-provisioner to get a necessary PVC:

    ❯ kubectl apply -f https://github.com/kubernetes-sigs/agent-sandbox/releases/download/v0.5.6/sandbox.yaml
    ❯ kubectl apply -f https://raw.githubusercontent.com/rancher/local-path-provisioner/v0.0.33/deploy/local-path-storage.yaml
    ❯ kubectl patch storageclass local-path -p '{"metadata": {"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'
    

    (the latest url for agent-standbox did not work, so I put found the latest version string and used that)

    And I installed openshell with the combined topology

    ❯ CHART=0.0.110
    ❯ helm upgrade --install openshell \
      oci://ghcr.io/nvidia/openshell/helm-chart \
      --version "$CHART" \
      --namespace openshell \
      --set supervisor.topology=combined \
      --set server.disableTls=true \
      --set server.auth.allowUnauthenticatedUsers=true
    Release "openshell" does not exist. Installing it now.
    Pulled: ghcr.io/nvidia/openshell/helm-chart:0.0.110
    Digest: sha256:8bf947365a46184e979a30039a7f22fe57e150827ab699bc688e9260f7c4af12
    NAME: openshell
    LAST DEPLOYED: Fri Aug 21 17:58:05 2026
    NAMESPACE: openshell
    STATUS: deployed
    REVISION: 1
    DESCRIPTION: Install complete
    TEST SUITE: None
    

    Which seems to have been applied

    ❯ kubectl -n openshell get cm -o yaml | grep -n -A2 -E 'topology'
    
    70:      topology                     = "combined"
    

    Then, trying to run a pod, the pod started just fine for me, baring some hiccup with slow network downloading sandbox pod; the cli timeouted, but the sandbox did come up

    ❯ kubectl -n openshell port-forward svc/openshell 8080:8080                                                                      
    
    Forwarding from 127.0.0.1:8080 -> 8080
    Forwarding from [::1]:8080 -> 8080
    Handling connection for 8080
    Handling connection for 8080
    Handling connection for 8080
    
    ❯ openshell gateway add http://127.0.0.1:8080 --local --name k8s
    ✓ Gateway 'k8s' added and set as active
      Endpoint: http://127.0.0.1:8080
      Type: local
      Auth: plaintext
    
    ❯ openshell sandbox create --name cap-setpcap-repro -- true
    
    
    Created sandbox: cap-setpcap-repro
    
    ✗ sandbox provisioning timed out after 300s. Last reported status: DependenciesNotReady: Pod exists with phase: Pending
    ✓ Sandbox allocated (4s)
    ✓ Image pulled (13 MB) (6s)                                                                                                                                                                                                                 Error:   × sandbox provisioning timed out after 300s. Last reported status: DependenciesNotReady: Pod exists with phase: Pending
    
    ❯ openshell sandbox create --name cap-setpcap-repro -- true
    
    Error:   × sandbox 'cap-setpcap-repro' already exists
      │ 
      │ hint: delete it first with: openshell sandbox delete <name>
      │       or use a different name
    
    ❯ kubectl -n openshell get sandboxes
    
    NAME                         READY   REASON              AGE
    default--cap-setpcap-repro   True    DependenciesReady   12m
    
    ❯ kubectl get pod -n openshell -o yaml default--cap-setpcap-repro
    apiVersion: v1
    kind: Pod
    metadata:
      annotations:
        agents.x-k8s.io/propagated-annotations: openshell.io/sandbox-id
        openshell.io/sandbox-id: fb863998-78c4-49e8-bbee-b944471864e2
      creationTimestamp: "2026-08-21T16:04:15Z"
      generation: 1
      labels:
        agents.x-k8s.io/sandbox-name-hash: 1fd73180
      name: default--cap-setpcap-repro
      namespace: openshell
      ownerReferences:
      - apiVersion: agents.x-k8s.io/v1beta1
        blockOwnerDeletion: true
        controller: true
        kind: Sandbox
        name: default--cap-setpcap-repro
        uid: 820af084-1316-41c8-87da-aa984ca61923
      resourceVersion: "3942"
      uid: c402b511-f1e5-4c08-8337-2f35b76d3683
    spec:
      automountServiceAccountToken: false
      containers:
      - command:
        - /opt/openshell/bin/openshell-sandbox
        - --workdir
        - /sandbox
        env:
        - name: OPENSHELL_SANDBOX_ID
          value: fb863998-78c4-49e8-bbee-b944471864e2
        - name: OPENSHELL_SANDBOX
          value: cap-setpcap-repro
        - name: OPENSHELL_ENDPOINT
          value: http://openshell.openshell.svc.cluster.local:8080
        - name: OPENSHELL_SANDBOX_COMMAND
          value: sleep infinity
        - name: OPENSHELL_TELEMETRY_ENABLED
          value: "true"
        - name: OPENSHELL_SSH_SOCKET_PATH
          value: /run/openshell/ssh.sock
        - name: OPENSHELL_K8S_SA_TOKEN_FILE
          value: /var/run/secrets/openshell/token
        - name: OPENSHELL_OCI_IMAGE_USER
        - name: OPENSHELL_SANDBOX_UID
          value: "1000"
        - name: OPENSHELL_SANDBOX_GID
          value: "1000"
        image: ghcr.io/nvidia/openshell-community/sandboxes/base:latest
        imagePullPolicy: Always
        name: agent
        resources: {}
        securityContext:
          appArmorProfile:
            type: Unconfined
          capabilities:
            add:
            - SYS_ADMIN
            - NET_ADMIN
            - SYS_PTRACE
            - SYSLOG
          runAsUser: 0
        terminationMessagePath: /dev/termination-log
        terminationMessagePolicy: File
        volumeMounts:
        - mountPath: /var/run/secrets/openshell
          name: openshell-sa-token
          readOnly: true
        - mountPath: /opt/openshell/bin
          name: openshell-supervisor-bin
          readOnly: true
        - mountPath: /sandbox
          name: workspace
      dnsPolicy: ClusterFirst
      enableServiceLinks: true
      initContainers:
      - command:
        - sh
        - -c
        - if [ ! -f /workspace-pvc/.workspace-initialized ]; then if [ -d /sandbox ];
          then tmp=$(mktemp) && rm -f "$tmp" && (cd /sandbox && find . -mindepth 1 -maxdepth
          1 -exec tar -cf "$tmp" {} +) && if [ -f "$tmp" ]; then tar -C /workspace-pvc
          --no-same-owner --no-same-permissions --touch -xf "$tmp" && rm -f "$tmp"; fi;
          fi && touch /workspace-pvc/.workspace-initialized; fi
        image: ghcr.io/nvidia/openshell-community/sandboxes/base:latest
        imagePullPolicy: Always
        name: workspace-init
        resources: {}
        securityContext:
          runAsUser: 0
        terminationMessagePath: /dev/termination-log
        terminationMessagePolicy: File
        volumeMounts:
        - mountPath: /workspace-pvc
          name: workspace
      nodeName: lima-k8s
      preemptionPolicy: PreemptLowerPriority
      priority: 0
      restartPolicy: Always
      schedulerName: default-scheduler
      securityContext:
        fsGroup: 1000
      serviceAccount: openshell-sandbox
      serviceAccountName: openshell-sandbox
      terminationGracePeriodSeconds: 30
      tolerations:
      - effect: NoExecute
        key: node.kubernetes.io/not-ready
        operator: Exists
        tolerationSeconds: 300
      - effect: NoExecute
        key: node.kubernetes.io/unreachable
        operator: Exists
        tolerationSeconds: 300
      volumes:
      - name: openshell-sa-token
        projected:
          defaultMode: 256
          sources:
          - serviceAccountToken:
              audience: openshell-gateway
              expirationSeconds: 3600
              path: token
      - image:
          pullPolicy: IfNotPresent
          reference: ghcr.io/nvidia/openshell/supervisor:7909fb5d0f54a06e26eb79e47885d7dd105aef24
        name: openshell-supervisor-bin
      - name: workspace
        persistentVolumeClaim:
          claimName: workspace-default--cap-setpcap-repro
    status:
      conditions:
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:04:19Z"
        observedGeneration: 1
        status: "True"
        type: PodReadyToStartContainers
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:13:14Z"
        observedGeneration: 1
        status: "True"
        type: Initialized
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:13:15Z"
        observedGeneration: 1
        status: "True"
        type: Ready
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:13:15Z"
        observedGeneration: 1
        status: "True"
        type: ContainersReady
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:04:19Z"
        observedGeneration: 1
        status: "True"
        type: PodScheduled
      containerStatuses:
      - containerID: containerd://cf81decc4cd667dfdf3515868c75e869899cb0a22386e3335982a9af4ad2b6f0
        image: ghcr.io/nvidia/openshell-community/sandboxes/base:latest
        imageID: ghcr.io/nvidia/openshell-community/sandboxes/base@sha256:aeef1c63f00e2913ea002ccb3aaf925f338b5c5d70e63576f0d95c16a138044e
        lastState: {}
        name: agent
        ready: true
        resources: {}
        restartCount: 0
        started: true
        state:
          running:
            startedAt: "2026-08-21T16:13:15Z"
        user:
          linux:
            gid: 0
            supplementalGroups:
            - 0
            - 1000
            uid: 0
        volumeMounts:
        - mountPath: /var/run/secrets/openshell
          name: openshell-sa-token
          readOnly: true
          recursiveReadOnly: Disabled
        - mountPath: /opt/openshell/bin
          name: openshell-supervisor-bin
          readOnly: true
          recursiveReadOnly: Disabled
        - mountPath: /sandbox
          name: workspace
      hostIP: 192.168.5.15
      hostIPs:
      - ip: 192.168.5.15
      initContainerStatuses:
      - containerID: containerd://00438f907a9c9c556fe6e921b3f41554f7f4e9d47c0a045e9543c3b27dc08d9e
        image: ghcr.io/nvidia/openshell-community/sandboxes/base:latest
        imageID: ghcr.io/nvidia/openshell-community/sandboxes/base@sha256:aeef1c63f00e2913ea002ccb3aaf925f338b5c5d70e63576f0d95c16a138044e
        lastState: {}
        name: workspace-init
        ready: true
        resources: {}
        restartCount: 0
        started: false
        state:
          terminated:
            containerID: containerd://00438f907a9c9c556fe6e921b3f41554f7f4e9d47c0a045e9543c3b27dc08d9e
            exitCode: 0
            finishedAt: "2026-08-21T16:13:13Z"
            reason: Completed
            startedAt: "2026-08-21T16:13:13Z"
        user:
          linux:
            gid: 0
            supplementalGroups:
            - 0
            - 1000
            uid: 0
        volumeMounts:
        - mountPath: /workspace-pvc
          name: workspace
      observedGeneration: 1
      phase: Running
      podIP: 10.244.0.10
      podIPs:
      - ip: 10.244.0.10
      qosClass: BestEffort
      resources: {}
      startTime: "2026-08-21T16:04:19Z"
    

    The above includes

        securityContext:
          appArmorProfile:
            type: Unconfined
          capabilities:
            add:
            - SYS_ADMIN
            - NET_ADMIN
            - SYS_PTRACE
            - SYSLOG
          runAsUser: 0
    

    For the openshell-0 pod

    ❯ kubectl get pod -n openshell -o yaml openshell-0
    apiVersion: v1
    kind: Pod
    metadata:
      annotations:
        checksum/gateway-config: 8d5fbcf2a0e531b1f7b2b998f6879a261b6284d0bef6c1bf2c44ee5c6d3897f8
      creationTimestamp: "2026-08-21T15:58:29Z"
      generateName: openshell-
      generation: 1
      labels:
        app.kubernetes.io/instance: openshell
        app.kubernetes.io/managed-by: Helm
        app.kubernetes.io/name: openshell
        app.kubernetes.io/version: 0.0.110
        apps.kubernetes.io/pod-index: "0"
        controller-revision-hash: openshell-5974c7bcbd
        helm.sh/chart: helm-chart-0.0.110
        statefulset.kubernetes.io/pod-name: openshell-0
      name: openshell-0
      namespace: openshell
      ownerReferences:
      - apiVersion: apps/v1
        blockOwnerDeletion: true
        controller: true
        kind: StatefulSet
        name: openshell
        uid: 76429d46-5f79-4592-a7bc-43832691d866
      resourceVersion: "2765"
      uid: 2c4953ec-585f-4934-9283-af1859516b07
    spec:
      containers:
      - args:
        - --config
        - /etc/openshell/gateway.toml
        - --db-url
        - sqlite:/var/openshell/openshell.db
        env:
        - name: OPENSHELL_GATEWAY_CREDENTIAL_KEY_ENCRYPTION_KEY
          valueFrom:
            secretKeyRef:
              key: key-encryption-key
              name: openshell-credential-storage-key-encryption-key
        - name: OPENSHELL_TELEMETRY_ENABLED
          value: "true"
        image: ghcr.io/nvidia/openshell/gateway:0.0.110
        imagePullPolicy: IfNotPresent
        livenessProbe:
          failureThreshold: 3
          httpGet:
            path: /healthz
            port: health
            scheme: HTTP
          initialDelaySeconds: 2
          periodSeconds: 5
          successThreshold: 1
          timeoutSeconds: 1
        name: openshell-gateway
        ports:
        - containerPort: 8080
          name: grpc
          protocol: TCP
        - containerPort: 8081
          name: health
          protocol: TCP
        - containerPort: 9090
          name: metrics
          protocol: TCP
        readinessProbe:
          failureThreshold: 3
          httpGet:
            path: /readyz
            port: health
            scheme: HTTP
          initialDelaySeconds: 1
          periodSeconds: 2
          successThreshold: 1
          timeoutSeconds: 1
        resources: {}
        securityContext:
          allowPrivilegeEscalation: false
          capabilities:
            drop:
            - ALL
          runAsNonRoot: true
          runAsUser: 1000
        startupProbe:
          failureThreshold: 30
          httpGet:
            path: /healthz
            port: health
            scheme: HTTP
          periodSeconds: 2
          successThreshold: 1
          timeoutSeconds: 1
        terminationMessagePath: /dev/termination-log
        terminationMessagePolicy: File
        volumeMounts:
        - mountPath: /var/openshell
          name: openshell-data
        - mountPath: /etc/openshell
          name: gateway-config
          readOnly: true
        - mountPath: /etc/openshell-jwt
          name: sandbox-jwt
          readOnly: true
        - mountPath: /var/run/secrets/kubernetes.io/serviceaccount
          name: kube-api-access-zrv4r
          readOnly: true
      dnsPolicy: ClusterFirst
      enableServiceLinks: true
      hostname: openshell-0
      nodeName: lima-k8s
      preemptionPolicy: PreemptLowerPriority
      priority: 0
      restartPolicy: Always
      schedulerName: default-scheduler
      securityContext:
        fsGroup: 1000
      serviceAccount: openshell
      serviceAccountName: openshell
      subdomain: openshell
      terminationGracePeriodSeconds: 5
      tolerations:
      - effect: NoExecute
        key: node.kubernetes.io/not-ready
        operator: Exists
        tolerationSeconds: 300
      - effect: NoExecute
        key: node.kubernetes.io/unreachable
        operator: Exists
        tolerationSeconds: 300
      volumes:
      - name: openshell-data
        persistentVolumeClaim:
          claimName: openshell-data-openshell-0
      - configMap:
          defaultMode: 420
          name: openshell-config
        name: gateway-config
      - name: sandbox-jwt
        secret:
          defaultMode: 256
          secretName: openshell-jwt-keys
      - name: kube-api-access-zrv4r
        projected:
          defaultMode: 420
          sources:
          - serviceAccountToken:
              expirationSeconds: 3607
              path: token
          - configMap:
              items:
              - key: ca.crt
                path: ca.crt
              name: kube-root-ca.crt
          - downwardAPI:
              items:
              - fieldRef:
                  apiVersion: v1
                  fieldPath: metadata.namespace
                path: namespace
    status:
      conditions:
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:03:19Z"
        observedGeneration: 1
        status: "True"
        type: PodReadyToStartContainers
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:03:18Z"
        observedGeneration: 1
        status: "True"
        type: Initialized
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:03:21Z"
        observedGeneration: 1
        status: "True"
        type: Ready
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:03:21Z"
        observedGeneration: 1
        status: "True"
        type: ContainersReady
      - lastProbeTime: null
        lastTransitionTime: "2026-08-21T16:03:18Z"
        observedGeneration: 1
        status: "True"
        type: PodScheduled
      containerStatuses:
      - containerID: containerd://50aa7e3ca3af8162adc186b7a6d078b63216c2ae6fc32aa60af2f232e7e1aa4e
        image: ghcr.io/nvidia/openshell/gateway:0.0.110
        imageID: ghcr.io/nvidia/openshell/gateway@sha256:398bf373bc8cf23bd9b1b1096bf5ecda814ff459c2f335ecc9c27aacdcef9fa1
        lastState: {}
        name: openshell-gateway
        ready: true
        resources: {}
        restartCount: 0
        started: true
        state:
          running:
            startedAt: "2026-08-21T16:03:19Z"
        user:
          linux:
            gid: 0
            supplementalGroups:
            - 0
            - 1000
            uid: 1000
        volumeMounts:
        - mountPath: /var/openshell
          name: openshell-data
        - mountPath: /etc/openshell
          name: gateway-config
          readOnly: true
          recursiveReadOnly: Disabled
        - mountPath: /etc/openshell-jwt
          name: sandbox-jwt
          readOnly: true
          recursiveReadOnly: Disabled
        - mountPath: /var/run/secrets/kubernetes.io/serviceaccount
          name: kube-api-access-zrv4r
          readOnly: true
          recursiveReadOnly: Disabled
      hostIP: 192.168.5.15
      hostIPs:
      - ip: 192.168.5.15
      observedGeneration: 1
      phase: Running
      podIP: 10.244.0.8
      podIPs:
      - ip: 10.244.0.8
      qosClass: BestEffort
      resources: {}
      startTime: "2026-08-21T16:03:18Z"
    

    That includes

        securityContext:
          allowPrivilegeEscalation: false
          capabilities:
            drop:
            - ALL
          runAsNonRoot: true
          runAsUser: 1000
    

    I see no crashloops or any errors besides seemingly innocuous

    ❯ kubectl logs -n openshell default--cap-setpcap-repro  -c agent
    2026-08-21T16:13:16.071Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/usr/local/bin/opencode container_path=/proc/38/root/usr/local/bin/opencode pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE.
    2026-08-21T16:13:16.072Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/usr/local/bin/opencode container_path=/proc/38/root/usr/local/bin/opencode pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE.
    2026-08-21T16:13:16.072Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/usr/bin/wget container_path=/proc/38/root/usr/bin/wget pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE.
    2026-08-21T16:13:16.079Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/usr/bin/wget container_path=/proc/38/root/usr/bin/wget pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE.
    2026-08-21T16:13:16.079Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/app/.venv/bin/python container_path=/proc/38/root/app/.venv/bin/python pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE.
    2026-08-21T16:13:16.079Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/app/.venv/bin/python3 container_path=/proc/38/root/app/.venv/bin/python3 pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE.
    2026-08-21T16:13:16.079Z WARN openshell_supervisor_network::opa: Cannot access container filesystem for symlink resolution: path=/app/.venv/bin/pip container_path=/proc/38/root/app/.venv/bin/pip pid=38 error=No such file or directory (os error 2). Binary paths in policy will be matched literally. If this binary is a symlink (e.g., /usr/bin/python3 -> python3.11), use the canonical path instead, or run with CAP_SYS_PTRACE.
    
  2. johntmyers commented on Sep 17, 2026

    @johntmyers
    Collaborator

    📋 triage-agent

    Closing as obsolete after RFC 0012. The failure is specific to the retired Kubernetes combined topology, where the privileged supervisor ran in the workload container. Kubernetes now uses separate supervisor and workload pods through #2942/#3144, so adding CAP_SETPCAP to that removed topology would restore privilege to a design OpenShell no longer uses.

  3. removed
    state:triage-neededOpened without agent diagnostics and needs triage
    on Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions