Skip to content

bug: sandbox SSH rejects loopback reverse port forwarding #4379

Description

@cjagwani

User Story

I use OpenShell to run an agent inside a sandbox while keeping its authenticated MCP service on the host loopback. I directly encountered the sandbox SSH server rejecting the loopback reverse forward, so the agent cannot start with its required host service available.

Problem Statement

OpenShell-generated SSH access supports session and direct TCP channels, but the embedded server rejects the standard SSH tcpip-forward request used by ssh -R. The failure is deterministic for both an occupied application port and unused high ports. Authentication succeeds, then OpenSSH exits because the server refuses the remote listener.

Impact / Why This Matters

A sandboxed agent cannot securely reach a host-only loopback service through its existing OpenShell SSH session. In the observed deployment this prevents the agent from becoming ready after a host reboot. Retrying, changing ports, and confirming that no sandbox process owns the port do not help. Exposing the host service on a non-loopback interface would weaken the boundary and is not an acceptable workaround.

Acceptance Criteria

  • A client can request -R 127.0.0.1:<sandbox-port>:127.0.0.1:<host-port> through an OpenShell-generated SSH config.
  • A sandbox connection to the forwarded loopback port round-trips bytes to the client-side loopback service.
  • Non-loopback bind addresses and invalid ports fail closed.
  • Duplicate binds fail without affecting an existing forward.
  • cancel-tcpip-forward and SSH disconnect close the listener and release the port.
  • The behavior has an automated regression test and survives sandbox/host restart validation.

Reproduction Steps

  1. Start an OpenShell sandbox and generate its supported SSH config:
    openshell sandbox ssh-config <sandbox> --workspace default > /tmp/openshell-ssh.conf
  2. Request a loopback-only reverse forward on an unused high port:
    ssh -F /tmp/openshell-ssh.conf -T \
      -o ExitOnForwardFailure=yes \
      -R 127.0.0.1:39991:127.0.0.1:39991 \
      openshell-<sandbox>.default true
  3. Observe authentication succeed followed by:
    Error: remote port forwarding failed for listen port 39991
    
  4. Repeat with other unused ports. Every tcpip-forward request is rejected.

Environment

  • OpenShell: 0.0.117-dev.268+g490055b42
  • OS: Ubuntu 24.04
  • Runtime, deployment, or integration: Direct Docker-backed sandbox, OpenShell-generated SSH config, loopback-only reverse forward

Logs

Authenticated to openshell-<sandbox>.default (via proxy) using "none".
Remote connections from 127.0.0.1:39991 forwarded to local address 127.0.0.1:39991
remote forward failure for: listen 127.0.0.1:39991, connect 127.0.0.1:39991
Error: remote port forwarding failed for listen port 39991

Activity

  1. cjagwani commented on Oct 9, 2026

    @cjagwani
    ContributorAuthor

    🏗️ build-plan

    Implement loopback-only SSH reverse forwarding in the sandbox SSH server.

    • Add a boundary-loopback listener contract so the listener is created inside the sandbox boundary rather than on the supervisor host.
    • Support tcpip-forward and cancel-tcpip-forward for loopback addresses only; reject wildcard, non-loopback, and invalid ports.
    • Track listeners per SSH connection so duplicate binds fail without disturbing the original, cancellation releases the port, and disconnect aborts all listeners.
    • Relay accepted sockets through forwarded-tcpip channels with bounded cleanup and structured SSH audit events.
    • Add regression coverage for byte round trips, invalid requests, duplicate binds, cancellation, and disconnect cleanup.
    • Run focused Rust checks plus the sandbox E2E path, then validate the exact failure on the isolated Gen2 VM without changing Stable or shared staging.

    The issue is still labeled state:triage-needed; implementation is proceeding from the direct operator request, without changing maintainer disposition.

  2. O96a commented on Oct 9, 2026

    @O96a

    That's sshd config, not the client — a loopback -R is gated server-side. The tell is whether -L still works: if local forwards are fine but -R gets refused, the sandbox almost certainly has AllowTcpForwarding local (or no), which is the usual hardening that kills remote forwards while leaving local ones alone.

    On OpenSSH >= 8.9 also check PermitListen — if it's set, it has to cover the exact 127.0.0.1:<port> you're asking the server to bind. GatewayPorts is a red herring here; it only governs binding to non-loopback addresses, so it won't explain a rejected loopback bind.

    The server log line is usually refused local port forward / administratively prohibited, and the client shows remote port forwarding failed for listen port N. If you can edit the sandbox sshd:

    AllowTcpForwarding yes
    PermitListen 127.0.0.1:*
    

    If the sshd is baked into the image, I'd stop fighting it and use OpenShell's own exec/port-forward channel to reach the host service instead of tunneling over SSH. A hardened sandbox sshd is a deliberate boundary, and the reverse-forward path is the wrong lever for it.

  3. cjagwani commented on Oct 9, 2026

    @cjagwani
    ContributorAuthor

    Thanks — that is the right distinction for OpenSSH sshd. This reproduction uses the embedded russh server in openshell-supervisor-process, so there is no sshd_config: -L works because channel_open_direct_tcpip is implemented, while -R is rejected because tcpip_forward and cancel_tcpip_forward are not implemented and russh defaults them to false. The existing OpenShell connector is supervisor-to-boundary; this use case needs boundary-to-client loopback. The patch keeps the hardening equivalent of PermitListen 127.0.0.1:* by accepting only loopback binds and rejecting wildcard or non-loopback addresses.

  4. johntmyers commented on Oct 9, 2026

    @johntmyers
    Collaborator

    @cjagwani this is a pretty narrow use case for opening up what could become a footgun in terms of allowing sandbox escape. have you explored running the MCP server in something that is addressable to the supervisor but doesn't require 0.0.0.0 binding on the actual host? for example in a container on the same network then you could allow access via a provider? I'm assuming by wanting to reverse forward to 127.0.0.1 on the host that's the scenario you're trying to avoid (binding MCP to 0.0.0.0?)

    I wouldn't consider reverse SSH connections to be a good durable connection for something that is critical to the agent's operations, that's why we moved away from using port forwarding via SSH as a means to host ingress traffic to listening servers in the agent sandbox

  5. cjagwani commented on Oct 9, 2026

    @cjagwani
    ContributorAuthor

    Thanks — the concrete path here is ACP, not MCP. Omnid owns a per-turn ACP adapter on host loopback, while Hermes inside the sandbox must dial that adapter for the duration of one turn. Gen2 therefore requests -R 127.0.0.1:<ephemeral>:127.0.0.1:<ephemeral>: neither side binds 0.0.0.0, the listener exists only for that SSH session, and the trusted host client—not sandbox code—requests the forward. OpenShell's existing local-forward path is the opposite direction (host client to sandbox), so it cannot carry this callback.

    I share the durability concern. If a first-class session-scoped reverse connector is the preferred architecture, I can pivot to that rather than expose generic -R. The required invariant is narrowly boundary-loopback to one client-owned loopback endpoint, with cancellation/disconnect cleanup. The current patch enforces the bind-side subset of that invariant: loopback only, wildcard/non-loopback rejected, and every listener leased to the SSH session.

  6. cjagwani commented on Oct 9, 2026

    @cjagwani
    ContributorAuthor

    Correction and E2E update after tracing the full launch path:

    • ACP itself still runs over SSH stdio.
    • The loopback ports passed to the launcher are per-turn HTTP MCP collaboration endpoints, so the proposed -R support is forwarding MCP traffic, not ACP.
    • I validated the supported host.openshell.internal alternative from a real managed sandbox: an explicitly allowed fixed gateway route reached the host-loopback collaboration listener.
    • The downstream managed deployment currently admits only signed, fixed gateway routes and does not yet carry the per-turn collaboration bearer through that route. Ad hoc provider/profile changes are correctly rejected.

    So the host-gateway approach is viable, but it needs a coordinated downstream policy/auth contract before it replaces reverse forwarding end to end. I am keeping #4384 open as the working fallback until that supported path passes a real task launch; I am not treating #4391's Docker E2E alone as closure for this issue.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions