Skip to content

feat: remove managed inference routes in favor of providers v2 #3172

Description

@krishicks

User Story

As an OpenShell operator, I want profile-backed providers to be the single way that sandboxes receive model-provider access, so that provider credentials, endpoint policy, and sandbox attachment lifecycle are configured through one consistent workflow.

Problem Statement

OpenShell currently exposes https://inference.local inside every sandbox and manages its upstream through a separate workspace-level inference route. It also carries a second sandbox-system route intended for supervisor-internal platform inference. Providers v2 now supports authoritative provider profiles, per-sandbox provider attachment, provider-owned network policy, and endpoint-bound credential injection. Keeping these managed inference routes duplicates that provider model with gateway-managed route state, separate openshell inference configuration, model pinning, and route-bundle delivery.

Remove both inference.local and sandbox-system, including the managed inference configuration and bundle-delivery surface, and make profile-backed providers the supported path for model-provider access. No production code currently calls the sandbox-system in-process API; it exists only in its implementation and tests.

Impact / Why This Matters

Today users must understand and maintain two overlapping abstractions: provider records/attachments for direct provider access and a workspace-global inference route for inference.local. The latter applies one provider and model across all sandboxes in a workspace, while provider attachments are scoped to individual sandboxes. This creates ambiguous ownership, prevents provider attachment from being the complete access contract, constrains clients to the proxy's supported API patterns, and expands the CLI, API, policy, routing, documentation, and test surface.

Users can already attach a provider and call its native endpoint, but current documentation presents that path alongside inference.local and still auto-allows the virtual host. Maintaining both paths adds operational and security-review cost without a distinct sandbox use case now that provider profiles own endpoints, L7 policy, and credential binding.

Proposed Design

Make profile-backed providers the only user-facing mechanism for granting a sandbox access to an inference provider:

  1. Users create or select a provider profile and provider instance.
  2. Users attach that provider when creating a sandbox or with openshell sandbox provider attach.
  3. Code in the sandbox calls the provider's native endpoint and uses the provider-defined environment/config contract. OpenShell continues to enforce the profile-derived network policy and inject credentials only at authorized endpoints.
  4. Detaching the provider revokes its policy and credential access for that sandbox.

Stop exposing, resolving, documenting, or implicitly allowing inference.local for sandbox workloads. Remove sandbox-system and its unused in-process supervisor API as well. Remove the managed inference CLI commands, gRPC configuration and bundle-delivery API, persisted route state, and supervisor route refresh/partitioning behavior. Existing installations should receive clear upgrade guidance that maps inference.local configurations to provider attachment and native endpoint usage rather than silently preserving the old feature. If a concrete platform-level inference consumer is introduced later, it should get an explicit provider-backed design based on that consumer's identity, policy, credential, and lifecycle requirements rather than retaining a speculative shared route.

Acceptance Criteria

  • New and upgraded sandboxes do not resolve, trust, implicitly allow, or route requests to inference.local.
  • The gateway no longer persists or delivers inference.local or sandbox-system routes, and removes the managed inference route/bundle API surface.
  • The openshell inference set|get|update|delete command surface is removed.
  • Attaching an inference-capable provider grants only that sandbox the profile-defined native endpoints, policy, and endpoint-bound credentials; detaching it revokes that access according to the provider lifecycle contract.
  • Built-in inference provider profiles contain the endpoint, protocol, credential, and configuration metadata required for supported clients to call their native APIs without inference.local.
  • OpenAI-compatible, Anthropic-compatible, and other supported inference-provider paths have end-to-end coverage through provider attachment and native endpoints.
  • Existing inference.local route state has deterministic upgrade behavior and the release notes provide a concrete migration from inference-route configuration to provider attachment.
  • README, published documentation, tutorials, architecture docs, CLI help, and public/internal agent skills no longer present inference.local as a sandbox feature.
  • Protobuf comments, generated SDK surfaces, policy defaults, certificates/DNS handling, router configuration, logging, and tests no longer carry sandbox-facing inference.local assumptions.
  • The sandbox-system route, supervisor route cache, and in-process system_inference() API are removed, with tests confirming no platform functionality depends on them.
  • The replacement preserves OpenShell's credential non-disclosure, endpoint binding, policy enforcement, hot-refresh, and OCSF logging guarantees.

Alternatives Considered

Keep inference.local alongside providers v2. This preserves compatibility but leaves users and maintainers with two overlapping ways to grant model access and retains workspace-global behavior that conflicts with per-sandbox provider attachment.

Automatically create an inference.local route when an inference-capable provider is attached. This makes attachment more convenient but preserves the virtual endpoint, model-routing abstraction, and duplicate routing/credential path that this proposal removes.

Deprecate inference.local indefinitely. A short, documented migration window may be appropriate, but maintaining both mechanisms as supported features would defer rather than eliminate the ambiguity and maintenance burden.

Retain sandbox-system for future platform functions. The route was introduced for a possible embedded agent harness such as policy analysis, but there is no production caller today. Keeping it would preserve much of the managed inference configuration, bundle, route-cache, and router machinery for a speculative consumer whose actual authorization and provider-lifecycle requirements are not yet defined.

Agent Investigation

Checklist

  • I've reviewed existing issues and the architecture docs
  • This is a design proposal, not a "please build this" request

Activity

  1. added this to the OpenShell 0.1.0 milestone on Sep 3, 2026
  2. self-assigned this
    on Sep 4, 2026
  3. johntmyers commented on Sep 4, 2026

    @johntmyers
    Collaborator

    🏗️ build-plan

    Implementation Plan

    Issue type: refactor
    Complexity: High
    Confidence: Medium — the end state is clear, but removal crosses the runtime, control plane, public APIs, SDKs, tests, and documentation.

    Summary

    Remove the managed inference feature set: inference.local, sandbox-system, workspace-scoped inference routes, route-bundle delivery, the openshell inference command/API/SDK surfaces, and the sandbox-local request-shaping router. Attached Providers v2 become the only supported mechanism for inference access.

    This is an intentional breaking change before the public re-release. No mixed-version compatibility, deprecation shim, route conversion, or automatic provider attachment is required. Legacy route rows will be purged, and the upgrade procedure will require existing sandboxes/supervisors to be recreated.

    This issue lands before the separate removal of automatically available Providers v2 profiles. Those profiles may remain temporarily as scaffolding, but replacement tests in this issue must explicitly import profile fixtures to prove the final profile import → provider create → sandbox attach → native endpoint workflow. Keep inference_capable as informational metadata for potential future observability; it must not imply current routing or policy behavior.

    Expected effort is approximately 8–12 engineer-days, preferably delivered as three focused PRs: dependency/control-plane cleanup, data-plane removal and replacement tests, then SDK/docs/skills cleanup.

    Scope

    • proto/inference.proto, crates/openshell-server/src/inference.rs, crates/openshell-server/src/multiplex.rs, and server auth tables: remove managed inference RPCs, route resolution, bundle delivery, and authorization plumbing.
    • crates/openshell-core/src/grpc_client.rs, crates/openshell-core/src/metadata.rs, crates/openshell-server/src/grpc/workspace.rs, and new SQLite/PostgreSQL migrations: remove route clients/object metadata/cascade logic and delete persisted object_type = 'inference_route' rows.
    • crates/openshell-supervisor-network/src/{proxy.rs,inference_routes.rs,l7/inference.rs,run.rs} and crates/openshell-supervisor-network/tests/system_inference.rs: remove the privileged inference.local interception, custom HTTP parser/request matcher, route cache/refresh loop, system_inference() API, and associated OCSF inference-routing events.
    • crates/openshell-sandbox/src/{main.rs,lib.rs}: remove --inference-routes / OPENSHELL_INFERENCE_ROUTES standalone configuration and route wiring.
    • crates/openshell-router/** plus server/supervisor/package dependencies: delete the inference-only router and its provider-specific header, path, model, body, timeout, retry, probe, and streaming behavior.
    • crates/openshell-core/src/inference.rs, crates/openshell-providers/src/{lib.rs,discovery.rs}, and crates/openshell-server/src/grpc/provider.rs: remove the old hard-coded managed-inference provider registry after relocating or deleting the small Providers v2 helpers that still depend on it. Do not remove the Providers v2 YAML profiles in this issue.
    • crates/openshell-cli/src/{main.rs,run.rs,tls.rs}: remove openshell inference set|get|update|delete and inference-only client setup.
    • Python and Go SDKs, generators, examples, fakes, and generated protobuf outputs: remove the inference route client surface and regenerate. Verify the TypeScript generation closure remains clean.
    • scripts/keycloak-realm.json, OIDC tests, and API authorization docs: remove inference-specific scopes and method mappings.
    • README.md, published docs, architecture docs, examples, and public/internal skills: remove inference.local as a supported feature and document the native-provider UX and breaking upgrade.

    Implementation Steps

    1. Decouple Providers v2 from the old inference registry.

      • Inventory the remaining consumers of crates/openshell-core/src/inference.rs.
      • Move or rewrite only the still-required provider-type aliases and Vertex configuration constants in provider-owned code.
      • Remove the route-only built-in endpoint suppression in crates/openshell-server/src/grpc/provider.rs; custom endpoints are expressed by explicitly imported profiles instead.
      • Preserve inference_capable in the profile schema and document it as informational metadata only.
    2. Remove the managed inference control plane.

      • Delete proto/inference.proto and stop registering or calling the service.
      • Delete server route resolution, endpoint verification, secret-bundle materialization, route persistence, and workspace route cleanup.
      • Add paired SQLite and PostgreSQL migrations deleting legacy inference_route object rows; leave historical migrations unchanged.
      • Remove inference-specific auth scopes and service-method authorization.
      • Remove the CLI command group.
    3. Remove the managed inference data plane.

      • Delete the pre-OPA CONNECT inference.local:443 branch and its implicit trust/policy behavior.
      • Delete route source loading, five-second refresh, user/system route partitioning, system_inference(), request-shape matching, model/path/body/header rewriting, and inference-specific streaming/probe behavior.
      • Remove --inference-routes from standalone sandbox/supervisor configuration.
      • Delete openshell-router after its last consumers are gone and clean Cargo/package references.
    4. Remove public SDK surfaces and regenerate.

      • Delete Python InferenceRouteClient exports/tests/generated modules.
      • Delete Go Inference() client interfaces, types, converters, fakes, examples, generated bindings, and codegen mappings.
      • Regenerate all affected protobuf outputs and verify no stale service descriptors remain.
    5. Replace tests with explicit imported-profile coverage.

      • Delete route CRUD, route bundle, router transformation, request-shape, sandbox-system, and inference.local success-path suites.
      • Add fixture profiles owned by the E2E suites; do not rely on automatically available profiles.
      • Exercise explicit profile import, provider creation, sandbox attachment, native endpoint access, credential binding, detach revocation, and the negative inference.local regression.
    6. Update documentation, architecture, examples, and skills.

      • Replace managed-route tutorials with explicit profile import/edit/create/attach and native endpoint guidance.
      • State that models, timeouts, native request shapes, SDK modes, and base URLs are application concerns.
      • State that the breaking upgrade purges managed routes and requires sandbox recreation.
      • Run the sync-agent-infra consistency workflow because this removes a crate and changes CLI/skill coverage.

    Test Plan

    • Unit tests

      • Update Providers v2 normalization/discovery tests after removing their dependency on the managed inference registry.
      • Remove obsolete router and route-resolution tests.
      • Update CLI parse/help tests to assert that the inference command group is absent.
      • Preserve tests establishing that inference_capable round-trips as metadata without affecting routing or policy.
    • Integration tests

      • Update server multiplex/auth/workspace tests so no inference RPC, descriptor, or route object remains.
      • Verify both database backends purge legacy inference_route rows while preserving provider records and attachments.
      • Remove Python/Go route-client tests and ensure the remaining SDKs compile without the deleted API.
    • E2E tests

      • Import an OpenAI-compatible fixture profile, create its provider, attach it to one sandbox, and call the fixture’s native endpoint.
      • Import an Anthropic-compatible fixture profile and exercise its native request shape.
      • Import a custom/self-hosted OpenAI-compatible fixture representing Ollama/LM Studio/vLLM-style users.
      • Verify unattached and detached sandboxes cannot use the profile’s endpoint-bound credential.
      • Verify raw credentials remain absent from sandbox-visible environment values, errors, and OCSF output.
      • Verify inference.local no longer resolves, receives special trust, bypasses ordinary policy, or routes requests.
      • Verify no platform functionality depends on sandbox-system.
    • Verification

      • mise run pre-commit
      • Relevant Rust, Python, and Go unit/integration suites
      • mise run test
      • mise run e2e:docker for the modified provider/sandbox E2E paths
      • mise run ci before the final PR

    Risks & Open Questions

    • The key hidden coupling is the old managed-inference registry being reused for Providers v2 aliases, Vertex constants, and endpoint activation. Decouple those consumers before deleting the module.
    • Do not compensate for inference.local removal by broadening credential scope. Native credential substitution must remain profile endpoint/path-bound and default-deny elsewhere.
    • API:INFERENCE model/token-aware events disappear with the request parser. Native traffic continues to emit ordinary network/HTTP and credential-binding OCSF events. inference_capable is retained so deeper observation can be designed separately later.
    • Automatically attaching the former route provider is explicitly out of scope; the old route was workspace-global, while provider attachment is a sandbox-scoped authorization decision.
    • Existing route model and timeout values are deleted rather than migrated. The user configures them in the native client.
    • Mixed-version operation is unsupported. Upgrade documentation must require replacing all running sandboxes/supervisors so no process retains a cached resolved route.
    • LSM compatibility: no new SELinux/AppArmor-specific behavior is expected because this removes interception code and adds no new /proc, exec, or process-identity behavior. Replacement E2E tests should use the existing provider/sandbox harness rather than add label-sensitive fork/exec probes.

    Documentation Impact

    • docs/reference/gateway-config.mdx: no expected change; this issue does not add, remove, or rename gateway TOML or compute-driver configuration.
    • docs/reference/sandbox-compute-drivers.mdx: no expected change unless implementation uncovers a driver-specific workaround.
    • Published docs: replace/remove docs/sandboxes/inference-routing.mdx, update provider profiles and Google Vertex guidance, rewrite local inference tutorials, and update policy, security, gateway-auth, how-it-works, supported-agents, and observability pages.
    • Architecture: remove managed inference from architecture/README.md, architecture/gateway.md, architecture/sandbox.md, architecture/security-policy.md, architecture/google-vertex-ai-provider.md, and relevant limits documentation.
    • Examples and skills: remove or rewrite examples/local-inference/, skills/debug-inference, skills/openshell-cli, policy-generation guidance, and the sync-agent-infra maintenance map.
    • Historical RFCs should remain historical; mark them superseded where useful rather than rewriting their original design record.

    UX Now / After

    Hosted provider

    Now

    openshell provider create \
      --name team-openai \
      --type openai \
      --credential OPENAI_API_KEY=<key>
    
    openshell inference set \
      --provider team-openai \
      --model gpt-4o \
      --timeout 120
    
    openshell sandbox create --name agent -- python app.py
    from openai import OpenAI
    
    client = OpenAI(
        base_url="https://inference.local/v1",
        api_key="unused",
    )
    client.chat.completions.create(
        model="anything",  # overwritten by the managed route
        messages=[{"role": "user", "content": "hi"}],
    )

    After

    # The test fixtures use this explicit flow immediately. Provider profiles
    # become examples-only in the subsequent built-in-profile removal issue.
    openshell provider profile import -f ./openai-native.yaml
    
    openshell provider create \
      --name team-openai \
      --type openai-native \
      --credential OPENAI_API_KEY=<key>
    
    openshell sandbox create \
      --name agent \
      --provider team-openai \
      -- python app.py
    import os
    from openai import OpenAI
    
    client = OpenAI(
        base_url="https://api.openai.com/v1",
        api_key=os.environ["OPENAI_API_KEY"],
        timeout=120,
    )
    client.chat.completions.create(
        model="gpt-4o",
        messages=[{"role": "user", "content": "hi"}],
    )

    The attachment is the complete access contract for this sandbox. The environment value remains an OpenShell placeholder and resolves only at profile-authorized endpoints.

    Local or custom OpenAI-compatible provider

    Now

    openshell provider create \
      --name ollama \
      --type openai \
      --credential OPENAI_API_KEY=unused \
      --config OPENAI_BASE_URL=http://host.openshell.internal:11434/v1
    
    openshell inference set --provider ollama --model qwen3.5:0.8b

    The application calls https://inference.local/v1; the router rewrites the request to the configured host endpoint and forces the model.

    After

    # The imported profile explicitly declares host.openshell.internal:11434,
    # its REST access rules, and the application binaries allowed to use it.
    openshell provider profile import -f ./ollama-openai.yaml
    openshell provider create --name ollama --type ollama-openai
    
    openshell sandbox create \
      --name local-agent \
      --provider ollama \
      --env OPENAI_BASE_URL=http://host.openshell.internal:11434/v1 \
      -- python app.py

    The application calls the native endpoint and supplies the real model. OpenShell enforces only the imported profile’s endpoint, path, binary, and credential contract.

    Upgrade behavior

    Before: one workspace route implicitly serves every sandbox.
    After:  each intended sandbox explicitly attaches its provider.
    
    Before: OpenShell owns endpoint, model, timeout, and request rewriting.
    After:  the profile owns allowed endpoint/credential policy;
            the application owns endpoint selection, model, timeout, and request shape.
    

    Existing managed route rows are not migrated. The supported breaking upgrade is: upgrade the gateway, purge route state through the schema migration, recreate all sandboxes, import/edit the required profiles, create providers, attach them to intended sandboxes, and update clients to native endpoints.


    Revision 1 — initial plan

  4. added
    state:acceptedA maintainer decided OpenShell should pursue this issue
    on Sep 4, 2026
  5. moved this from Todo to In progress in OpenShell Roadmapon Sep 4, 2026
  6. added a commit that references this issue on Sep 4, 2026
    6599bfe
  7. johntmyers commented on Sep 4, 2026

    @johntmyers
    Collaborator

    🏗️ build-from-issue-agent

    Implementation Complete

    PR: #3195

    What was built

    Removed the managed inference route control plane, inference.local proxy path, request-shape matching, built-in router crate, and inference SDK surfaces. Inference now uses explicitly imported provider profiles attached to sandboxes and provider-native endpoints; inference_capable remains informational.

    Tests

    • Unit/integration: full Rust, Python, Go, and TypeScript suites pass via mise run test and mise run ci
    • E2E: Docker Rust suite passes; Python Docker suite passes with 86 passed and 81 OIDC-only tests skipped

    Docs updated

    • Provider-backed inference guide with breaking migration and Now/After UX
    • Local inference, Ollama, LM Studio, Vertex, provider, security, observability, architecture, and troubleshooting documentation

    The issue will auto-close when the PR is merged.

  8. added 4 commits that reference this issue on Sep 4, 2026
    36b3abd
    8e6dbb3
    cd23991
    0e5d7dd
  9. moved this from In progress to Done in OpenShell Roadmapon Sep 9, 2026
  10. added a commit that references this issue on Sep 9, 2026
    f4dc6be
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

state:acceptedA maintainer decided OpenShell should pursue this issuestate:agent-readyApproved for agent implementation

Type

No type

Projects

Relationships

None yet

Development

No branches or pull requests

Issue actions