惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Security Latest
Security Latest
G
Google Developers Blog
量子位
WordPress大学
WordPress大学
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
大猫的无限游戏
大猫的无限游戏
Apple Machine Learning Research
Apple Machine Learning Research
阮一峰的网络日志
阮一峰的网络日志
S
SegmentFault 最新的问题
Blog — PlanetScale
Blog — PlanetScale
B
Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
P
Proofpoint News Feed
美团技术团队
V
Visual Studio Blog
Last Week in AI
Last Week in AI
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
F
Fortinet All Blogs
博客园 - Franky
The Register - Security
The Register - Security
O
OpenAI News
Google DeepMind News
Google DeepMind News
A
Arctic Wolf
罗磊的独立博客
博客园 - 叶小钗
博客园 - 聂微东
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Y
Y Combinator Blog
SecWiki News
SecWiki News
T
Tor Project blog
月光博客
月光博客
S
Secure Thoughts
博客园 - 【当耐特】
Help Net Security
Help Net Security
D
Docker
Recent Announcements
Recent Announcements
GbyAI
GbyAI
B
Blog RSS Feed
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
T
The Blog of Author Tim Ferriss
Webroot Blog
Webroot Blog
V
Vulnerabilities – Threatpost
Forbes - Security
Forbes - Security
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Hacker News: Ask HN
Hacker News: Ask HN
Cyberwarzone
Cyberwarzone
宝玉的分享
宝玉的分享
Cisco Talos Blog
Cisco Talos Blog
I
InfoQ
Microsoft Security Blog
Microsoft Security Blog

Sealos Blog

Build a Full-Stack App with Claude Code + InsForge — Zero Backend Code | Sealos Blog InsForge vs Supabase: Which Backend for AI-Powered Development? | Sealos Blog Claude Code Metrics Dashboard: Grafana Setup (2026) | Sealos Blog What Is RustFS? Apache 2.0 MinIO Alternative (2026) | Sealos Blog Claude Code Mobile: iPhone, Android & SSH (2026) | Sealos Blog Eaglercraft Server Hosting: Fast Setup (2026) | Sealos Blog An Honest Review: Migrating a Complex Microservice App from Heroku to Sealos | Sealos Blog The Ultimate Guide to Kubernetes Audit Logging for Security and Compliance | Sealos Blog Cost Optimization Shootout: Sealos Autonomous FinOps vs. Kubecost Manual Reports | Sealos Blog For CTOs: How to Cut Your Cloud Bill by 50% Without Sacrificing Performance | Sealos Blog Building Resilient Systems: A Deep Dive into Sealos High-Availability and Auto-Failover | Sealos Blog Building a Scalable Event-Driven Architecture with Sealos Managed Kafka | Sealos Blog Beyond kubectl apply: 5 GitOps Best Practices for Production-Ready CI/CD on Sealos | Sealos Blog Advanced RAG Pipelines: Why Your Choice of Vector Database (like Milvus) Matters | Sealos Blog Advanced MLOps: How to Monitor and Evaluate LLM Applications in Production | Sealos Blog A Developer's Guide to Kubernetes RBAC: Securing Your Cluster the Easy Way with Sealos | Sealos Blog A CISO's Guide to Cloud Development: Securing the CI/CD Pipeline with Sealos DevBox | Sealos Blog What is Kubernetes Multi-Tenancy? A Guide for Platform Engineers | Sealos Blog What is Infrastructure from Code (IfC)? The Next Step After Infrastructure as Code (IaC) | Sealos Blog What is GitOps? A Beginner's Guide to "Push-to-Deploy" Workflows | Sealos Blog What is eBPF? The Future of Kubernetes Networking and Security | Sealos Blog What is an "AI-Native" Platform? (And Why You Need One for MLOps) | Sealos Blog What is an Agentic Workflow? Building the Next Generation of AI Apps | Sealos Blog What is a Kubernetes Chargeback Model (And How Does it Save You Money?) | Sealos Blog What is a "Headless" Development Environment? (And How it Works with VS Code) | Sealos Blog What is a Graph-Based Vector Database? (And When to Use It Over Milvus) | Sealos Blog What is a "Cloud Operating System"? The Next Evolution of PaaS Explained | Sealos Blog The Real Cost of EKS: How Sealos Delivers a Simpler, Cheaper Kubernetes Experience | Sealos Blog The 3 Types of Kubernetes Autoscaling (HPA, VPA, CA) and How Sealos Manages Them for You | Sealos Blog Sealos vs Vercel: Why a Cloud OS Beats a Frontend Platform for Full-Stack Apps | Sealos Blog Sealos vs. Render vs. Fly.io: A 2025 Guide to the Best Heroku Alternatives | Sealos Blog Sealos vs. OpenShift: Kubernetes for Developers vs. Kubernetes for Ops Teams | Sealos Blog Sealos vs. Netlify: When to Choose a Full Kubernetes Platform over a Static Site Hoster | Sealos Blog Sealos vs. DigitalOcean App Platform: A Head-to-Head Comparison on Cost, Features, and Scalability | Sealos Blog Sealos vs. AWS Elastic Beanstalk: The Modern PaaS for Developers Who Hate YAML | Sealos Blog Sealos DevBox vs. AWS Cloud9: Why Your CDE Should Be Platform-Agnostic | Sealos Blog For Developers: Stop Wasting Time on DevOps. A 10-Minute Guide to Shipping Faster with DevBox. | Sealos Blog Deploying n8n with Docker: From Local Setups to a Radically Simple Cloud Alternative | Sealos Blog The Impact of Prompt Bloat: How the Sealos AI Proxy Can Cache Queries and Cut LLM Costs | Sealos Blog The FinOps Playbook: How to Implement Kubernetes Chargebacks and Showbacks with Sealos | Sealos Blog Smoke Testing for ML Pipelines: Catching Data and Model Errors Before They Hit Production | Sealos Blog Optimizing PostgreSQL Performance: A Guide to Sealos Managed Database Tuning | Sealos Blog Managing Kubernetes Multi-Tenancy: How Sealos Enforces Resource Quotas and Network Policies | Sealos Blog From Days to Minutes: How to Standardize Developer Environments for Your Entire Engineering Org | Sealos Blog For Platform Engineers: How to Build a Golden Path IDP (Internal Developer Platform) with Sealos | Sealos Blog For FinOps Managers: The 5 Leakiest Buckets in Your Kubernetes Budget (And How to Plug Them) | Sealos Blog For Educators & IT Admins: How to Provide a Secure, Scalable Cloud Lab for 1000+ Students on a Budget | Sealos Blog What is a Vector Database? A Beginner's Guide to Milvus, Pinecone, and More | Sealos Blog Why Your Microservices Architecture is Failing (And How a Cloud OS Can Fix It) | Sealos Blog The Power of Autoscaling: A Deep Dive into HPA, VPA, and Cluster Autoscaler | Sealos Blog The Total Economic Impact of Cloud Development Environments (CDEs) | Sealos Blog The Illustrated Guide to the Kubernetes Control Plane | Sealos Blog The MLOps Lifecycle Explained: From Data Prep to Model Deployment | Sealos Blog Beyond Vercel's AI Cloud: The Case for an AI-Native Operating System | Sealos Blog The Architecture of a Modern AI Application: A 2025 Blueprint | Sealos Blog GitHub Codespaces is Great, But Your Workflow is Incomplete. Here's Why. | Sealos Blog The Best Heroku Alternatives in 2025 for Scalability and Cost | Sealos Blog CAST AI vs. Kubecost vs. Sealos: Choosing the Right K8s Cost Management Tool | Sealos Blog DevBox vs. Gitpod vs. Replit: An Unbiased Comparison for 2025 | Sealos Blog Unlocking Hidden Savings: A Guide to Using Spot Instances Safely in Kubernetes | Sealos Blog Can a CDE Really Replace Your MacBook Pro? A Performance Benchmark | Sealos Blog The End of "Works on My Machine": Achieving 100% Reproducible Builds with DevBox | Sealos Blog The Ultimate Guide to GPU Provisioning and Management in Kubernetes | Sealos Blog Rightsizing Kubernetes Workloads: How to Stop Wasting Money on CPU and Memory Requests | Sealos Blog The 2025 Guide to Kubernetes Cost Optimization: 10 Strategies to Cut Your Bill in Half | Sealos Blog FinOps for Startups: How to Build a Cost-Conscious Culture from Day One | Sealos Blog How to Onboard a New Developer in Under 5 Minutes with Sealos DevBox | Sealos Blog Calculating Kubernetes Costs: A Breakdown of EKS, GKE, and AKS Pricing Models | Sealos Blog Case Study: How We Reduced Our Kubernetes Bill by 87% with Sealos | Sealos Blog Are You Overpaying for Managed Kubernetes? The True Cost of Vendor Lock-in | Sealos Blog Beyond Monitoring: How Sealos Autonomously Optimizes Your Cloud Spend | Sealos Blog A Practical Guide to Kubernetes Security: Hardening Your Cluster in 2025 | Sealos Blog A Secure-by-Design Development Workflow with Isolated Cloud Environments | Sealos Blog Setting Up a Collaborative Python Data Science Environment with DevBox | Sealos Blog Using the Sealos AI Proxy to Manage and Cache LLM API Calls | Sealos Blog Migration Guide: Moving Your Node.js & Postgres App from Heroku to Sealos in Under an Hour | Sealos Blog Serving Machine Learning Models at Scale: A Guide to Inference Optimization | Sealos Blog Headless Development with Sealos: Using Your Local VS Code with a Powerful Cloud Backend | Sealos Blog How to Build and Deploy a RAG Pipeline with Llama 3 and Milvus on Sealos | Sealos Blog From Localhost to Production in 15 Minutes: A Full-Stack CDE Workflow with Sealos DevBox | Sealos Blog GitOps on Autopilot: Implementing a CI/CD Pipeline with Sealos and GitHub Actions | Sealos Blog Fine-Tuning Open-Source LLMs on a Budget with Sealos | Sealos Blog From Docker Compose to Kubernetes: A Simple Migration Path with Sealos | Sealos Blog Building an AI Agentic Workflow with LangChain and Sealos | Sealos Blog What is Helm for Kubernetes? The Ultimate Package Manager Explained | Sealos Blog What is a Custom Resource Definition (CRD) in Kubernetes? | Sealos Blog What is a Kubernetes StatefulSet? A Practical Guide | Sealos Blog What is a Kubernetes Ingress Controller? A Guide to Smart Traffic Routing | Sealos Blog What is a Kubernetes Operator? Automating Complex Applications | Sealos Blog What is a Kubernetes Service? A Simple Guide for Developers | Sealos Blog Streamlining Your CI/CD Pipeline with a DevBox Build Environment | Sealos Blog Why Standardized Development Environments Are Key to Team Velocity | Sealos Blog What Is GitHub Codespace? | Sealos Blog DevBox Install? Skip It Entirely. Get a Ready-to-Code Environment in One Click with Sealos DevBox. | Sealos Blog How to Set Up a DevBox: The Ultimate Guide to 1-Click Cloud Development | Sealos Blog Empowering Indie Devs and Startup Teams: How Sealos DevBox Accelerates Agile Development | Sealos Blog From Chaos to Consistency: How Sealos DevBox Transforms Enterprise Development Workflows | Sealos Blog From Campus Labs to Cloud Freedom: How Sealos DevBox Supercharges Student Development | Sealos Blog How Sealos DevBox Cut Container Commit Time from 15 Minutes to 1 Second | Sealos Blog DevBox vs Codespaces: Which Remote Dev Environment Fits You Best? | Sealos Blog
Kubernetes NodePort Exhaustion: SSH Gateway Solution | Sealos Blog
Sealos · 2026-01-22 · via Sealos Blog

TL;DR: We built an SSH Gateway that routes all Devbox SSH traffic through port 2233, eliminating Kubernetes NodePort exhaustion. Two routing modes: (1) public key fingerprint lookup, (2) username-encoded destination with Agent Forwarding. Source code →


An SSH Gateway routes SSH connections to multiple backend servers through one port. In Kubernetes, it prevents NodePort exhaustion by multiplexing SSH traffic.

Instead of allocating one NodePort per SSH destination, a gateway multiplexes connections through a shared port. SSH uses zero NodePorts, and configurations stay stable across restarts.

What you'll learn:

  • Why one-port-per-Devbox SSH doesn't scale in Kubernetes
  • How SSH's three-layer design enables single-port routing
  • Two routing methods: public key fingerprint lookup and username encoding
  • How SSH Agent Forwarding authenticates users without exposing private keys
  • Implementation details with open-source code

We built this for Sealos Devbox, our cloud development environment where each workspace needs SSH access. The architecture applies to any multi-tenant Kubernetes platform facing port constraints.

It started with a support ticket.

A support ticket hit our queue:

I'm on the Starter plan. 4 NodePorts max. I have 3 Devboxes, and each one eats a port just for SSH. That leaves me one port for everything else. Normal shutdown won't release the port—it just sits there, holding my quota hostage. Cold shutdown frees it up, but when I restart, the port number changes and my SSH config breaks. What am I supposed to do?

We pulled the logs. Searched for similar complaints. One became ten. Ten became fifty.

Then enterprise teams started writing in. "We run 100 Devboxes for 50 engineers. Even on our plan, the port quota runs out fast. We forced cold shutdowns to reclaim ports—but reconfiguring SSH after every restart is killing our productivity."

Same story, over and over. We had forced users into a choice that should never have existed.

How Kubernetes NodePort Allocation Works

NodePort is a Kubernetes service type that exposes applications to external traffic. When you create a NodePort service, Kubernetes allocates a port from a fixed range (typically 30000–32767) and opens it on every node. Traffic arriving at that port gets routed to the correct Pod.

In our original architecture, every Devbox got its own NodePort for SSH access—port 30123 for Alice, port 30124 for Bob, and so on.

This approach doesn't scale. NodePorts are finite, and cloud platforms enforce quotas per tenant to prevent exhaustion.

On Sealos Cloud, each subscription tier has a cap:

PlanNodePort Limit
Starter4
Hobby8
Standard16
Pro32

The math doesn't work. One Devbox, one SSH port. Three Devboxes on a Starter plan uses 75% of your quota just to log in.

The Quota Trap: Normal Shutdown vs. Cold Shutdown

We introduced two shutdown modes to address port exhaustion. Each created its own problem.

Normal shutdown keeps the Devbox suspended but preserves its port assignment. Your SSH config stays valid, but the port still counts against your quota even while the Devbox sleeps. You're paying for resources you aren't using.

Cold shutdown releases the port back to the pool. Your quota is available again, but when you restart, Kubernetes assigns a new port number and your SSH config breaks. Every restart means reconfiguring your local ~/.ssh/config, updating scripts, and notifying teammates.

We'd turned an infrastructure constraint into a daily headache. Users shouldn't need to know what a NodePort is, or have to choose between hoarding quota and staying productive.

This wasn't a feature request. It was a trap we'd built ourselves.

We were stuck for a week. Every fix we tried hit the same wall: ports are finite, and we couldn't create more quota.

During a Friday review, our architect stopped us mid-sentence.

"Wait. Why does each Devbox need its own port? SSH can multiplex connections. One port should be enough."

Silence. Someone pulled up the RFC.

SSH (Secure Shell) has been around since 1995—Tatu Ylönen built it at Helsinki University of Technology to replace Telnet and other insecure protocols. We all use SSH every day. But reading the spec? That's different.

We started reading, and found our answer in the protocol's three-layer architecture.

SSH Transport Layer: Key Exchange and Host Authentication

The SSH Transport Layer is the protocol's foundation. It handles encryption negotiation and server identity verification before any user credentials are exchanged.

When you type:

...the transport layer kicks in immediately. Here's what happens:

  1. Protocol version exchange. Client and server agree on SSH-2.0.
  2. Algorithm negotiation. Both sides agree on encryption ciphers, MAC algorithms, and key exchange methods.
  3. Key exchange. The Diffie-Hellman exchange establishes a shared session key without transmitting it.
  4. Server authentication. The server sends its Host Key, a public key that uniquely identifies this server.

That last step caught our attention. You've seen this prompt:

Type yes, and the key gets saved to ~/.ssh/known_hosts:

If the Host Key changes unexpectedly on subsequent connections, SSH throws a warning. Possible man-in-the-middle attack.

This raised our first design constraint. If we ran multiple Gateway replicas behind a load balancer, each replica would generate its own Host Key. Users would hit different replicas on different connections and see that warning every time.

Solution: deterministic key generation. Every Gateway replica derives the same Ed25519 Host Key from a shared seed stored in Kubernetes Secrets. Same key, every replica, every time.

SSH Authentication Layer: Public Key Verification Flow

Once the transport layer establishes a secure channel, the SSH Authentication Layer proves the user's identity.

SSH supports several authentication methods—password, keyboard-interactive, public key. Sealos Devbox uses public key authentication.

Here's how it works:

  1. Client sends public key. The client offers its public key to the server.
  2. Server checks authorized_keys. If the key matches an entry in ~/.ssh/authorized_keys, the server proceeds.
  3. Challenge-response. The server generates a random challenge and asks the client to prove it has the private key.
  4. Client signs with private key. The client uses its private key to create a signature and sends it back.
  5. Server verifies. The server checks the signature against the public key. Match? Access granted.

The private key never crosses the wire. The server verifies possession without ever seeing the secret.

We stared at this flow. Then it clicked.

The server sees the public key during authentication. Not after—during. That's our routing hook.

If our Gateway intercepts the authentication handshake, we can extract the public key, look it up in a database, and determine which Devbox the user wants—before the connection is fully established.

One key, one Devbox. No port numbers. No user configuration.

SSH Connection Layer: Channel Multiplexing Explained

After authentication, the SSH Connection Layer takes over. This is where channel multiplexing happens.

A single TCP connection can carry multiple independent channels—separate logical data streams that don't interfere with each other.

You've used this without realizing. Open an SSH session, run a command, transfer a file with scp, forward a port with -L—all over one connection, each in its own channel.

The most common channel type is the session channel, which provides your interactive shell. When you run ls -la:

  1. Client requests a new session channel (type: "session")
  2. Server allocates the channel and confirms
  3. Client requests a pseudo-terminal (PTY) for interactive use
  4. Client sends the "exec" or "shell" request
  5. Command output flows back through the channel
  6. Channel closes when the command completes or the shell exits

Each channel manages its own flow control using a sliding window protocol. Multiple channels operate in parallel without collision.

That's when the architecture clicked.

We'd been thinking about this wrong. We didn't need to route users at the port level. We could route them at the channel level.

  • Authentication tells us who (which public key, which user)
  • The connection layer handles where (multiplexing to the right Devbox)

One port, many Devboxes, no collision.

The protocol had the answer all along. We just needed to build the Gateway.

A traditional SSH server is a destination. You connect to it, and it runs your commands locally.

Our SSH Gateway is a relay. You connect to it, and it forwards you somewhere else—your actual Devbox, running in a different Pod, in a different namespace, with a different IP.

From the outside, the topology looks similar to before. One connection, one Devbox. But the resource accounting is different.

Old architecture: One NodePort per Devbox. A hundred Devboxes for your team? A hundred ports consumed—before you expose a single real service.

New architecture: One NodePort. Port 2233. Every Devbox, every user, every connection—all through that single entry point.

SSH no longer competes for quota. Those 4 or 8 or 32 NodePorts in your plan? They're yours again—for databases, APIs, web servers, the things that need external exposure.

But we'd traded one problem for another.

When everyone enters through the same door, how do you know where each person is headed? A thousand users hit port 2233 simultaneously. Which Devbox does each one want?

We found two answers, each rooted in a different layer of the SSH protocol.

Routing by Public Key Fingerprint

The simplest approach: use the SSH key itself as the routing identifier.

Every Devbox already has a key pair. When you create a Devbox, we automatically generate an Ed25519 key pair, which is faster and more compact than RSA. The public key stays on the Devbox in ~/.ssh/authorized_keys. The private key gets downloaded to your local machine.

One key, one Devbox. The mapping already exists. We just had to use it.

The Gateway maintains a lookup table:

When a user connects, the Gateway intercepts the authentication handshake, extracts the offered public key, computes its fingerprint, and queries the routing table.

  1. User initiates SSH connection to gateway.sealos.io:2233
  2. Transport layer completes (encryption established)
  3. Authentication begins—client offers its public key
  4. Gateway extracts the public key before forwarding
  5. Gateway queries: "Which Devbox owns this key?"
  6. Match found → Gateway opens a new SSH connection to that Devbox's internal IP
  7. Gateway relays all subsequent traffic between user and Devbox

No match? Connection rejected. The key isn't registered to any Devbox.

From the user's perspective, the experience works:

Shut down the Devbox. Restart it a week later. Cold boot it after a month of vacation. The key stays the same. The config stays the same. SSH works.

This approach works because of what we discovered in the authentication layer: the Gateway sees the public key during the handshake, before the connection is fully established. Each public key fingerprint is cryptographically unique, so there's no ambiguity. No collisions, no guessing.

For most Sealos users, this is the only mode they'll ever need. But not everyone wants to use our generated keys.

Routing by Username Encoding (bring your own key)

Some developers have their own SSH keys—keys they've used for years, backed up properly, integrated into their workflows. Some teams manage keys centrally through identity providers and don't want per-Devbox credentials scattered across employee laptops. Some security-conscious users rely on hardware tokens like YubiKey, where the private key is generated on-device and never leaves.

For these users, the public-key-as-identifier approach breaks down. The Gateway has never seen their key. There's nothing to look up.

We needed a different routing signal. Something explicit. Something the user controls.

The answer was obvious once we saw it: encode the destination in the SSH username.

Standard SSH usernames are strings. The protocol doesn't care what's in them. Nobody said we couldn't embed routing information.

We settled on a format: user@namespace-devbox

Instead of:

You write:

The Gateway parses ubuntu@ns-alice-my-workspace and extracts:

ComponentValue
Actual usernameubuntu
Namespacens-alice
Devbox namemy-workspace

Now the Gateway knows where to route—no key lookup required.

But authentication still needs to work. The Devbox's authorized_keys contains the user's personal public key. To complete the connection, the Gateway needs to authenticate to the Devbox on the user's behalf.

Here's the problem: the only credential that will unlock that door is the user's private key. And we don't have it. We can't have it.

Asking users to upload private keys to a third-party service? Non-starter. Security teams would reject it. Compliance frameworks would flag it. Users would refuse.

The Gateway sat in an impossible position: a valid connection from the user on one side, a locked Devbox door on the other, and no key for either.

We needed a way to borrow the user's signing capability without possessing their private key.

The answer had been in SSH all along. We just hadn't realized it.

Someone on the team spoke up during a late-night debugging session.

"What about Agent Forwarding?"

We'd all used it before. Bastion hosts. Jump boxes. Hopping between servers without copying private keys around. Standard ops practice. Been in OpenSSH since the early days.

But how did it actually work? None of us could explain the mechanics.

We pulled up the docs. And found exactly what we needed.

What Is SSH Agent and How It Works

You've probably used SSH Agent without thinking about it.

Every time you type ssh user@server and don't get prompted for your key's passphrase, that's the agent at work. It's a daemon running on your laptop, holding your private keys in memory, ready to sign authentication challenges.

The agent does two things:

  1. Stores your private keys (encrypted in memory, unlocked by passphrase)
  2. Signs data when asked (without revealing the key)

Here's what happens when you connect to a server:

The SSH client never touches your private key file directly. It asks the agent: "Sign this." The agent signs it.

The private key never leaves the agent's memory. Servers only see signatures—never the key itself.

What does this buy you? Even if a remote server is compromised, attackers can't steal your private key. They might hijack your agent connection while you're logged in and request signatures for their own purposes, but they can't extract the key for later use. The moment you disconnect, their access ends.

Hardware tokens like YubiKey take this further. The private key is generated inside the device and cannot be exported. The token acts as its own agent, requiring a physical tap for each signature.

How Agent Forwarding Relays Authentication

Your agent can sign things locally. But what happens when you're on a remote server and need to authenticate to another server?

You could copy your private key to the remote machine. Don't. That defeats the entire point.

SSH Agent Forwarding solves this. It extends your local agent's reach through the SSH connection, across networks, into remote servers without moving the key.

Or in your SSH config:

That -A flag (or ForwardAgent yes) tells the server: "I have an agent running locally. If you need signatures, ask me."

But how does a remote server talk to a daemon on your laptop? The mechanics are worth understanding, especially since our Gateway exploits them.

Step 1: Establish the Primary Connection

First, the normal SSH handshake:

Nothing special yet. Just a regular SSH connection with a flag set.

Step 2: Client Requests Agent Forwarding

Once the session is established, the client opens a session channel and sends a special request:

The server acknowledges the capability.

Step 3: Server Opens a Reverse Channel

When the server needs a signature, it opens a channel back to the client:

This is a reverse channel—initiated by the server, pointing back to the client. Most SSH traffic flows client-to-server. Agent forwarding reverses that.

On the server side, SSH creates a Unix domain socket (something like /tmp/ssh-XXXX/agent.12345) and sets SSH_AUTH_SOCK to point to it. Any process that writes to this socket is talking through the SSH connection, back to your laptop's agent.

Step 4: Sign Through the Tunnel

Suppose the server tries to SSH somewhere else—an internal database, a Git server, another jump host. The remote target demands authentication. The server doesn't have a key. But it has your forwarded agent.

The signature travels: Agent → Client → Server → Target. The key travels: nowhere.

This is the classic jump-host pattern. SSH into a bastion, forward your agent, hop to internal servers. Authenticate everywhere with your personal key—even though that key exists only on your laptop.

The Security Trade-off

Agent forwarding isn't free. While your connection is active, anyone with root on the server can hijack your forwarded agent. They can't steal the key—but they can use it. Request signatures. Authenticate as you. Access whatever your key unlocks.

The moment you disconnect, their access ends. But the window exists.

Mitigations:

  • Only forward to servers you trust
  • Use ssh-add -c to require confirmation for each signature
  • Prefer ProxyJump (-J) when possible—it avoids forwarding altogether

For our Gateway, the trade-off works. We control the infrastructure. The forwarding window is limited to each connection's duration, and the alternative—asking users to upload private keys—is worse.

Implementing Agent Forwarding in the Gateway

All the pieces were on the table. We just had to wire them together.

The Gateway sits in the middle of two SSH connections: one from the user, one to the Devbox. The user's connection carries an agent channel. The Devbox connection needs signatures. Our job: bridge the two.

Here's the complete flow for BYOK (Bring Your Own Key) routing:

Four phases, two connections, zero key exposure.

Phase 1: User connects. Gateway parses the encoded username to determine their destination. No key lookup needed—the destination is in the connection request.

Phase 2: Sets up the agent channel. The user's SSH client says "I can forward my agent." The Gateway accepts. A reverse channel opens, pointing back to the user's laptop.

Phase 3: The Gateway opens a new SSH connection to the Devbox. The Devbox demands authentication. The Gateway doesn't have that key. But it has a tunnel to someone who does.

The authentication challenge travels backward—Gateway to user, user to agent. The agent signs. The signature travels forward—agent to user, user to Gateway, Gateway to Devbox. Door opens.

Phase 4: Two session channels, one on each side, stitched together. Keystrokes flow in, output flows out. The Gateway becomes invisible.

The private key never touches the Gateway. The Gateway sees signatures—cryptographic proof that the key exists—but the key itself stays locked in the user's agent.

Hardware tokens work without modification. YubiKey, smart cards, TPM-backed credentials—if your agent supports it, this flow works too. The signing happens on your device. Tap to confirm.

Gateway Configuration Example

For BYOK mode, users configure their SSH client once:

Breaking down each directive:

DirectivePurpose
HostName gateway.sealos.ioConnect to the Gateway, not the Devbox directly
Port 2233The Gateway's single shared port
User ubuntu@ns-alice-my-workspaceEncodes both the target user (ubuntu) and routing info (ns-alice, my-workspace)
IdentityFile ~/.ssh/id_ed25519Your personal key (also in the Devbox's authorized_keys)
ForwardAgent yesEnable agent forwarding so the Gateway can authenticate to the Devbox

After this one-time setup:

Cold shutdown the Devbox, restart it next week, spin up new Devboxes in different namespaces. The Gateway config never changes. Add a new entry to ~/.ssh/config, and you're done.

For teams managing many Devboxes, the pattern scales:

Same key, same config structure, different destinations. The Gateway handles routing, the user's agent handles authentication.

Remember the support ticket that started this?

I'm on the Starter plan. 4 NodePorts max. I have 3 Devboxes, and each one eats a port just for SSH. That leaves me one port for everything else.

NodePorts consumed by SSH now: zero.

The Gateway runs on port 2233. That's one port for the entire cluster, not one per Devbox or per user. A hundred developers with a hundred Devboxes each? Still one port.

Those 4 NodePorts on the Starter plan? They're available again for databases, web servers, APIs—services that need external exposure.

"Cold shutdown frees up the port, but when I restart, the port number changes and my SSH config breaks."

SSH configuration changes after restart: zero.

The Gateway's address never changes. The routing happens by key fingerprint or username encoding, not by port number. Restart your Devbox after a month of vacation. Your ~/.ssh/config still works. Your IDE's remote development extension still connects. Your deployment scripts still run.

The choice that shouldn't have existed (quota versus productivity) is gone.

What we shipped

The SSH Gateway is now the default for all Sealos Devbox SSH access. Both routing modes are production-ready:

ModeBest ForConfiguration
Key fingerprint routingMost users; simple setupUse the auto-generated Devbox key
Username encoding + Agent ForwardingBYOK users, hardware tokens, team-managed keysEncode destination in username, enable ForwardAgent

The Gateway handles thousands of concurrent connections, routing them through a single port to Devboxes across the cluster.

Open source

We didn't build proprietary magic. The Gateway combines SSH protocol fundamentals (transport, authentication, connection multiplexing, agent forwarding) with routing logic.

The source code is available:

github.com/labring/sealos/tree/main/service/sshgate

If you're building a cloud development platform, a multi-tenant Kubernetes environment, or any system where SSH port sprawl is becoming a problem, the code is available to fork, deploy, or learn from.

We didn't add a feature. We removed a trap by reading the protocol spec and using what SSH already provided.


Further Reading:

FAQ