Repository navigation
feat: remove managed inference routes in favor of providers v2 #3172
Description
Activity
- added a parent issue
on Sep 3, 2026 🏗️ build-plan
Implementation Plan
Issue type:
refactor
Complexity: High
Confidence: Medium — the end state is clear, but removal crosses the runtime, control plane, public APIs, SDKs, tests, and documentation.Summary
Remove the managed inference feature set:
inference.local,sandbox-system, workspace-scoped inference routes, route-bundle delivery, theopenshell inferencecommand/API/SDK surfaces, and the sandbox-local request-shaping router. Attached Providers v2 become the only supported mechanism for inference access.This is an intentional breaking change before the public re-release. No mixed-version compatibility, deprecation shim, route conversion, or automatic provider attachment is required. Legacy route rows will be purged, and the upgrade procedure will require existing sandboxes/supervisors to be recreated.
This issue lands before the separate removal of automatically available Providers v2 profiles. Those profiles may remain temporarily as scaffolding, but replacement tests in this issue must explicitly import profile fixtures to prove the final
profile import → provider create → sandbox attach → native endpointworkflow. Keepinference_capableas informational metadata for potential future observability; it must not imply current routing or policy behavior.Expected effort is approximately 8–12 engineer-days, preferably delivered as three focused PRs: dependency/control-plane cleanup, data-plane removal and replacement tests, then SDK/docs/skills cleanup.
Scope
proto/inference.proto,crates/openshell-server/src/inference.rs,crates/openshell-server/src/multiplex.rs, and server auth tables: remove managed inference RPCs, route resolution, bundle delivery, and authorization plumbing.crates/openshell-core/src/grpc_client.rs,crates/openshell-core/src/metadata.rs,crates/openshell-server/src/grpc/workspace.rs, and new SQLite/PostgreSQL migrations: remove route clients/object metadata/cascade logic and delete persistedobject_type = 'inference_route'rows.crates/openshell-supervisor-network/src/{proxy.rs,inference_routes.rs,l7/inference.rs,run.rs}andcrates/openshell-supervisor-network/tests/system_inference.rs: remove the privilegedinference.localinterception, custom HTTP parser/request matcher, route cache/refresh loop,system_inference()API, and associated OCSF inference-routing events.crates/openshell-sandbox/src/{main.rs,lib.rs}: remove--inference-routes/OPENSHELL_INFERENCE_ROUTESstandalone configuration and route wiring.crates/openshell-router/**plus server/supervisor/package dependencies: delete the inference-only router and its provider-specific header, path, model, body, timeout, retry, probe, and streaming behavior.crates/openshell-core/src/inference.rs,crates/openshell-providers/src/{lib.rs,discovery.rs}, andcrates/openshell-server/src/grpc/provider.rs: remove the old hard-coded managed-inference provider registry after relocating or deleting the small Providers v2 helpers that still depend on it. Do not remove the Providers v2 YAML profiles in this issue.crates/openshell-cli/src/{main.rs,run.rs,tls.rs}: removeopenshell inference set|get|update|deleteand inference-only client setup.- Python and Go SDKs, generators, examples, fakes, and generated protobuf outputs: remove the inference route client surface and regenerate. Verify the TypeScript generation closure remains clean.
scripts/keycloak-realm.json, OIDC tests, and API authorization docs: remove inference-specific scopes and method mappings.README.md, published docs, architecture docs, examples, and public/internal skills: removeinference.localas a supported feature and document the native-provider UX and breaking upgrade.
Implementation Steps
-
Decouple Providers v2 from the old inference registry.
- Inventory the remaining consumers of
crates/openshell-core/src/inference.rs. - Move or rewrite only the still-required provider-type aliases and Vertex configuration constants in provider-owned code.
- Remove the route-only built-in endpoint suppression in
crates/openshell-server/src/grpc/provider.rs; custom endpoints are expressed by explicitly imported profiles instead. - Preserve
inference_capablein the profile schema and document it as informational metadata only.
- Inventory the remaining consumers of
-
Remove the managed inference control plane.
- Delete
proto/inference.protoand stop registering or calling the service. - Delete server route resolution, endpoint verification, secret-bundle materialization, route persistence, and workspace route cleanup.
- Add paired SQLite and PostgreSQL migrations deleting legacy
inference_routeobject rows; leave historical migrations unchanged. - Remove inference-specific auth scopes and service-method authorization.
- Remove the CLI command group.
- Delete
-
Remove the managed inference data plane.
- Delete the pre-OPA
CONNECT inference.local:443branch and its implicit trust/policy behavior. - Delete route source loading, five-second refresh, user/system route partitioning,
system_inference(), request-shape matching, model/path/body/header rewriting, and inference-specific streaming/probe behavior. - Remove
--inference-routesfrom standalone sandbox/supervisor configuration. - Delete
openshell-routerafter its last consumers are gone and clean Cargo/package references.
- Delete the pre-OPA
-
Remove public SDK surfaces and regenerate.
- Delete Python
InferenceRouteClientexports/tests/generated modules. - Delete Go
Inference()client interfaces, types, converters, fakes, examples, generated bindings, and codegen mappings. - Regenerate all affected protobuf outputs and verify no stale service descriptors remain.
- Delete Python
-
Replace tests with explicit imported-profile coverage.
- Delete route CRUD, route bundle, router transformation, request-shape,
sandbox-system, andinference.localsuccess-path suites. - Add fixture profiles owned by the E2E suites; do not rely on automatically available profiles.
- Exercise explicit profile import, provider creation, sandbox attachment, native endpoint access, credential binding, detach revocation, and the negative
inference.localregression.
- Delete route CRUD, route bundle, router transformation, request-shape,
-
Update documentation, architecture, examples, and skills.
- Replace managed-route tutorials with explicit profile import/edit/create/attach and native endpoint guidance.
- State that models, timeouts, native request shapes, SDK modes, and base URLs are application concerns.
- State that the breaking upgrade purges managed routes and requires sandbox recreation.
- Run the
sync-agent-infraconsistency workflow because this removes a crate and changes CLI/skill coverage.
Test Plan
-
Unit tests
- Update Providers v2 normalization/discovery tests after removing their dependency on the managed inference registry.
- Remove obsolete router and route-resolution tests.
- Update CLI parse/help tests to assert that the
inferencecommand group is absent. - Preserve tests establishing that
inference_capableround-trips as metadata without affecting routing or policy.
-
Integration tests
- Update server multiplex/auth/workspace tests so no inference RPC, descriptor, or route object remains.
- Verify both database backends purge legacy
inference_routerows while preserving provider records and attachments. - Remove Python/Go route-client tests and ensure the remaining SDKs compile without the deleted API.
-
E2E tests
- Import an OpenAI-compatible fixture profile, create its provider, attach it to one sandbox, and call the fixture’s native endpoint.
- Import an Anthropic-compatible fixture profile and exercise its native request shape.
- Import a custom/self-hosted OpenAI-compatible fixture representing Ollama/LM Studio/vLLM-style users.
- Verify unattached and detached sandboxes cannot use the profile’s endpoint-bound credential.
- Verify raw credentials remain absent from sandbox-visible environment values, errors, and OCSF output.
- Verify
inference.localno longer resolves, receives special trust, bypasses ordinary policy, or routes requests. - Verify no platform functionality depends on
sandbox-system.
-
Verification
mise run pre-commit- Relevant Rust, Python, and Go unit/integration suites
mise run testmise run e2e:dockerfor the modified provider/sandbox E2E pathsmise run cibefore the final PR
Risks & Open Questions
- The key hidden coupling is the old managed-inference registry being reused for Providers v2 aliases, Vertex constants, and endpoint activation. Decouple those consumers before deleting the module.
- Do not compensate for
inference.localremoval by broadening credential scope. Native credential substitution must remain profile endpoint/path-bound and default-deny elsewhere. API:INFERENCEmodel/token-aware events disappear with the request parser. Native traffic continues to emit ordinary network/HTTP and credential-binding OCSF events.inference_capableis retained so deeper observation can be designed separately later.- Automatically attaching the former route provider is explicitly out of scope; the old route was workspace-global, while provider attachment is a sandbox-scoped authorization decision.
- Existing route model and timeout values are deleted rather than migrated. The user configures them in the native client.
- Mixed-version operation is unsupported. Upgrade documentation must require replacing all running sandboxes/supervisors so no process retains a cached resolved route.
- LSM compatibility: no new SELinux/AppArmor-specific behavior is expected because this removes interception code and adds no new
/proc, exec, or process-identity behavior. Replacement E2E tests should use the existing provider/sandbox harness rather than add label-sensitive fork/exec probes.
Documentation Impact
docs/reference/gateway-config.mdx: no expected change; this issue does not add, remove, or rename gateway TOML or compute-driver configuration.docs/reference/sandbox-compute-drivers.mdx: no expected change unless implementation uncovers a driver-specific workaround.- Published docs: replace/remove
docs/sandboxes/inference-routing.mdx, update provider profiles and Google Vertex guidance, rewrite local inference tutorials, and update policy, security, gateway-auth, how-it-works, supported-agents, and observability pages. - Architecture: remove managed inference from
architecture/README.md,architecture/gateway.md,architecture/sandbox.md,architecture/security-policy.md,architecture/google-vertex-ai-provider.md, and relevant limits documentation. - Examples and skills: remove or rewrite
examples/local-inference/,skills/debug-inference,skills/openshell-cli, policy-generation guidance, and thesync-agent-inframaintenance map. - Historical RFCs should remain historical; mark them superseded where useful rather than rewriting their original design record.
UX Now / After
Hosted provider
Now
openshell provider create \ --name team-openai \ --type openai \ --credential OPENAI_API_KEY=<key> openshell inference set \ --provider team-openai \ --model gpt-4o \ --timeout 120 openshell sandbox create --name agent -- python app.py
from openai import OpenAI client = OpenAI( base_url="https://inference.local/v1", api_key="unused", ) client.chat.completions.create( model="anything", # overwritten by the managed route messages=[{"role": "user", "content": "hi"}], )
After
# The test fixtures use this explicit flow immediately. Provider profiles # become examples-only in the subsequent built-in-profile removal issue. openshell provider profile import -f ./openai-native.yaml openshell provider create \ --name team-openai \ --type openai-native \ --credential OPENAI_API_KEY=<key> openshell sandbox create \ --name agent \ --provider team-openai \ -- python app.py
import os from openai import OpenAI client = OpenAI( base_url="https://api.openai.com/v1", api_key=os.environ["OPENAI_API_KEY"], timeout=120, ) client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "hi"}], )
The attachment is the complete access contract for this sandbox. The environment value remains an OpenShell placeholder and resolves only at profile-authorized endpoints.
Local or custom OpenAI-compatible provider
Now
openshell provider create \ --name ollama \ --type openai \ --credential OPENAI_API_KEY=unused \ --config OPENAI_BASE_URL=http://host.openshell.internal:11434/v1 openshell inference set --provider ollama --model qwen3.5:0.8bThe application calls
https://inference.local/v1; the router rewrites the request to the configured host endpoint and forces the model.After
# The imported profile explicitly declares host.openshell.internal:11434, # its REST access rules, and the application binaries allowed to use it. openshell provider profile import -f ./ollama-openai.yaml openshell provider create --name ollama --type ollama-openai openshell sandbox create \ --name local-agent \ --provider ollama \ --env OPENAI_BASE_URL=http://host.openshell.internal:11434/v1 \ -- python app.py
The application calls the native endpoint and supplies the real model. OpenShell enforces only the imported profile’s endpoint, path, binary, and credential contract.
Upgrade behavior
Before: one workspace route implicitly serves every sandbox. After: each intended sandbox explicitly attaches its provider. Before: OpenShell owns endpoint, model, timeout, and request rewriting. After: the profile owns allowed endpoint/credential policy; the application owns endpoint selection, model, timeout, and request shape.Existing managed route rows are not migrated. The supported breaking upgrade is: upgrade the gateway, purge route state through the schema migration, recreate all sandboxes, import/edit the required profiles, create providers, attach them to intended sandboxes, and update clients to native endpoints.
Revision 1 — initial plan
- addedstate:agent-readyApproved for agent implementationApproved for agent implementationstate:acceptedA maintainer decided OpenShell should pursue this issueA maintainer decided OpenShell should pursue this issue
on Sep 4, 2026 - added a commit that references this issue
on Sep 4, 2026 🏗️ build-from-issue-agent
Implementation Complete
PR: #3195
What was built
Removed the managed inference route control plane,
inference.localproxy path, request-shape matching, built-in router crate, and inference SDK surfaces. Inference now uses explicitly imported provider profiles attached to sandboxes and provider-native endpoints;inference_capableremains informational.Tests
- Unit/integration: full Rust, Python, Go, and TypeScript suites pass via
mise run testandmise run ci - E2E: Docker Rust suite passes; Python Docker suite passes with 86 passed and 81 OIDC-only tests skipped
Docs updated
- Provider-backed inference guide with breaking migration and Now/After UX
- Local inference, Ollama, LM Studio, Vertex, provider, security, observability, architecture, and troubleshooting documentation
The issue will auto-close when the PR is merged.
- Unit/integration: full Rust, Python, Go, and TypeScript suites pass via
- added 4 commits that reference this issue
on Sep 4, 2026 - added a commit that references this issue
on Sep 9, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
User Story
As an OpenShell operator, I want profile-backed providers to be the single way that sandboxes receive model-provider access, so that provider credentials, endpoint policy, and sandbox attachment lifecycle are configured through one consistent workflow.
Problem Statement
OpenShell currently exposes
https://inference.localinside every sandbox and manages its upstream through a separate workspace-level inference route. It also carries a secondsandbox-systemroute intended for supervisor-internal platform inference. Providers v2 now supports authoritative provider profiles, per-sandbox provider attachment, provider-owned network policy, and endpoint-bound credential injection. Keeping these managed inference routes duplicates that provider model with gateway-managed route state, separateopenshell inferenceconfiguration, model pinning, and route-bundle delivery.Remove both
inference.localandsandbox-system, including the managed inference configuration and bundle-delivery surface, and make profile-backed providers the supported path for model-provider access. No production code currently calls thesandbox-systemin-process API; it exists only in its implementation and tests.Impact / Why This Matters
Today users must understand and maintain two overlapping abstractions: provider records/attachments for direct provider access and a workspace-global inference route for
inference.local. The latter applies one provider and model across all sandboxes in a workspace, while provider attachments are scoped to individual sandboxes. This creates ambiguous ownership, prevents provider attachment from being the complete access contract, constrains clients to the proxy's supported API patterns, and expands the CLI, API, policy, routing, documentation, and test surface.Users can already attach a provider and call its native endpoint, but current documentation presents that path alongside
inference.localand still auto-allows the virtual host. Maintaining both paths adds operational and security-review cost without a distinct sandbox use case now that provider profiles own endpoints, L7 policy, and credential binding.Proposed Design
Make profile-backed providers the only user-facing mechanism for granting a sandbox access to an inference provider:
openshell sandbox provider attach.Stop exposing, resolving, documenting, or implicitly allowing
inference.localfor sandbox workloads. Removesandbox-systemand its unused in-process supervisor API as well. Remove the managed inference CLI commands, gRPC configuration and bundle-delivery API, persisted route state, and supervisor route refresh/partitioning behavior. Existing installations should receive clear upgrade guidance that mapsinference.localconfigurations to provider attachment and native endpoint usage rather than silently preserving the old feature. If a concrete platform-level inference consumer is introduced later, it should get an explicit provider-backed design based on that consumer's identity, policy, credential, and lifecycle requirements rather than retaining a speculative shared route.Acceptance Criteria
inference.local.inference.localorsandbox-systemroutes, and removes the managed inference route/bundle API surface.openshell inference set|get|update|deletecommand surface is removed.inference.local.inference.localroute state has deterministic upgrade behavior and the release notes provide a concrete migration from inference-route configuration to provider attachment.inference.localas a sandbox feature.inference.localassumptions.sandbox-systemroute, supervisor route cache, and in-processsystem_inference()API are removed, with tests confirming no platform functionality depends on them.Alternatives Considered
Keep
inference.localalongside providers v2. This preserves compatibility but leaves users and maintainers with two overlapping ways to grant model access and retains workspace-global behavior that conflicts with per-sandbox provider attachment.Automatically create an
inference.localroute when an inference-capable provider is attached. This makes attachment more convenient but preserves the virtual endpoint, model-routing abstraction, and duplicate routing/credential path that this proposal removes.Deprecate
inference.localindefinitely. A short, documented migration window may be appropriate, but maintaining both mechanisms as supported features would defer rather than eliminate the ambiguity and maintenance burden.Retain
sandbox-systemfor future platform functions. The route was introduced for a possible embedded agent harness such as policy analysis, but there is no production caller today. Keeping it would preserve much of the managed inference configuration, bundle, route-cache, and router machinery for a speculative consumer whose actual authorization and provider-lifecycle requirements are not yet defined.Agent Investigation
docs/providers/profiles.mdxdocuments per-sandbox attach/detach, just-in-time provider policy composition, and endpoint-bound credential injection. It still lists inference mounting from attached providers as a roadmap item and points users to the separateinference.localmodel.docs/sandboxes/inference-routing.mdxdocumentsinference.localas one workspace-level provider/model route shared by every sandbox, configured withopenshell inference.proto/inference.protoandcrates/openshell-server/src/inference.rsdefault empty route names toinference.localwhile also carrying the distinctsandbox-systemroute.sandbox-systemwas introduced by feat: add sandbox.inference.local endpoint for system-level inference #207/feat(inference): add sandbox-system inference route for platform-level inference #209 for possible platform functions such as an embedded policy-analysis harness. The currentsystem_inference()API has no production call sites; only its implementation and integration tests reference it.inference.localreferences across 54 files, including CLI/server/router/policy code, protobuf and generated SDKs, e2e tests, README, architecture docs, published docs, and public/internal skills.Checklist