Edgeweir
Guides

Layer-4 forwarding

TCP / UDP port forwarding: port pools, L4 apps, PROXY protocol, DNS, statistics, and node ports.

Concepts

TermDefinition
L4 appForwards TCP connections or UDP sessions on one port of every node of a cluster to its origins.
Port poolA range of ports the cluster's L4 apps may use, per protocol (TCP, UDP, TCP + UDP).
Reserved portA port of the cluster's HTTP / HTTPS listeners; never part of a port pool.
UDP sessionThe datagrams from one client address and port; one session until the idle timeout.
PROXY protocolA header at the start of a connection that carries the client's address and port; v1 is text, v2 binary.

How it works

  1. Port ranges are set on the cluster's Port pools tab; an L4 app's port must fall inside one.
  2. Saving an L4 app publishes the cluster's configuration revision (through the configuration canary as for sites; disabling and deleting reach every node at once); every node of the cluster listens on the port and forwards connections to the origins by weight.
  3. With DNS steering on for the cluster, every enabled app has the CNAME target <app UUID>.<cluster domain>, through which clients reach healthy nodes.
  4. Nodes report per app and per minute the connections, refusals, peak concurrency, and bytes, shown on the app's Analytics tab.

Nodes need the capability l4-v1, see Node requirements.

Set up port pools

  1. Open Clusters & nodes, select the cluster, and switch to the Port pools tab (/clusters?tab=ports, shown once the cluster has a node).
  2. Click Add port pool, select the Protocol (TCP, UDP, or TCP + UDP), and fill in First port and Last port. For a single port both are the same.
  3. Click Save. The console shows Saved.
  4. Open these ports per protocol in the node hosts' firewall and the cloud security group; container nodes also publish them, see Add nodes.

Reserved ports at the top of the card lists the ports of the cluster's HTTP / HTTPS listeners; L4 apps next to the title opens the cluster's L4 app list.

ItemBehavior
Ports1024–65535; the first port must not be above the last ("The first port must not be above the last")
CountAt most 64 port pools per cluster
OverlapPools of one protocol must not overlap, and TCP + UDP overlaps both TCP and UDP ("Port pools overlap: …"); the conflicting pools are marked
ShrinkingRefused when the pools no longer contain the port of an app, disabled apps included ("Ports in use by …")
PublishingPort pools only check app ports: saving publishes no configuration revision, nodes are not affected
Auditcluster.port_pools_update, with the pools before and after as metadata

Create an L4 app

  1. Open Sites, switch to L4 apps with the Sites | L4 apps switch above the list (/l4), and click New L4 app. Searching L4 apps with ⌘K / Ctrl+K opens the page too.

  2. Fill in Name (at most 100 characters); with several clusters, select the Cluster.

  3. Select the Protocol (TCP or UDP) and fill in Listen port; for a range of ports also Last port. The hint Port pools: … below lists the pools of that protocol; when the cluster has none, the hint reads "The cluster has no TCP port pool" and Set up port pools opens the cluster's Port pools tab.

  4. Under Origins, fill in Origin, Port, and Weight, and turn on Backup where needed; Add origin adds more, see Origins and health checks.

  5. Set PROXY protocol, Timeouts, Passive health check, IP lists, and Limits per node as needed.

  6. Enabled is on by default. Click Create; the console shows "Created, revision #N".

  7. Verify: the app appears in the list and CNAME target shows <app UUID>.<cluster domain>; connect to the port on any node:

    nc -vz <node IP> 9000                      # TCP: the connection succeeds
    dig @<node IP> -p 5353 example.com +short  # UDP, here forwarding to a DNS origin
  8. The app's Analytics tab starts showing connections (reported every minute, after about 1–2 minutes).

List and app page

The L4 apps list is sorted by port, with the columns Name, Protocol, Listen port, Origins, CNAME target, and Enabled; with several clusters, select All clusters or one cluster. Row menu: Edit, Analytics, Delete.

Click a name to open the app page (/l4/<app ID>): the Overview tab shows the status, listen port, cluster, CNAME target (with the target of every line), origins, PROXY protocol, timeouts, passive health check, IP lists, limits per node, and update time, with Edit and Delete; the Analytics tab is described in Statistics.

ActionWhereBehavior
EditRow menu or Edit on the app page, dialog Edit L4 appSame fields as for creating; the cluster cannot change. Saving shows "Saved, revision #N"
Disable / enableThe Enabled switch in the list, or the switch next to Status on the app page; asks first, and disabling notes "Nodes stop listening on the port and its DNS record is removed"A disabled app keeps its port and settings, is not shipped to nodes, and has no DNS record; setting the current state publishes nothing
DeleteRow menu or Delete on the app page, confirm "Delete L4 app {name}?"Deletes the app, its origins, and its statistics
ChangeRevision reasonAudit
CreateL4 application {app} createdl4_app.create
EditL4 application {app} updatedl4_app.update (changed fields with before and after)
Disable / enableL4 application {app} updatedl4_app.disable / l4_app.enable
DeleteL4 application {app} deletedl4_app.delete

Origins and health checks

FieldValuesDefaultEffect
OriginHost name or IP—Follows the origin address restrictions: special-purpose addresses only inside the origin allow list; nodes check every address a host name resolves to
Port1–65535—Origin port; may differ from the listen port
Weight1–1001Random choice by weight
BackupOn / offOffUsed only while every primary origin is down

An app has 1–32 origins, at least one of them not a backup ("Keep at least one origin that is not a backup").

ItemBehavior
ChoiceHealthy primary origins at random by weight; when every primary is down, only healthy backups; when all are down, all of them in turn, primaries first
RetriesThe next origin only after a failed connection to the origin (refused, timed out), at most 3 origins per connection; no switch once connected
Host namesResolved by the resolver nodes use for sites; a failed resolution, or only refused addresses, counts as one failure

Passive health check and timeouts:

FieldValuesDefaultEffect
Passive health check · Failures before down1–1003After this many failed connections (for UDP also an unreachable origin, such as ICMP port unreachable) the origin is taken out
Passive health check · Retry after (seconds)1–360030How long the origin stays out before it is tried again; also how long failures count, one success clears them
Timeouts · Connect (seconds)0.1–605Timeout for connecting to an origin
Timeouts · Idle (seconds)1–86400TCP 600, UDP 30How long without data in either direction ends a connection or session; switching the protocol in the dialog sets the protocol's default (unless edited)

Health state lives on each node only: it is not reported to the console and not shown in the UI. On the node host, list the origins currently down:

Node
sudo curl -s --unix-socket /run/edgeweir-node/control.sock http://localhost/v1/l4

down in the response lists {app_id, origin_id, down_until}.

Port ranges and origin ports

FieldValuesDefaultEffect
Last portEmpty, or above the listen port; at most 1000 ports per appEmptyThe app listens on every port from Listen port to Last port; the whole range must be inside the protocol's port pools (adjacent pools may share it) and must not overlap another app's port or range
Origin portFixed / Same as the portFixedFixed: each origin's own port; same: the origin port is the port the connection arrived on (9005 of range 9000-9010 connects to the origin's 9005), the origins' port fields stay empty
ItemRule
Cluster limitA cluster's L4 apps listen on at most 2048 ports together (ranges count each port; "A cluster's L4 applications use at most 2048 ports")
Node resourcesA TCP range takes one listening socket per port; UDP one per port and worker, see Ports and firewalls
Statistics and DNSPer app, as for a single port
Node capabilityRanges, origins on the arriving port and TLS termination need l4-v2

TLS termination

A TCP app may choose a TLS certificate: the node terminates TLS and connects to the origins over plain TCP.

FieldValuesDefaultEffect
TLS certificateNo TLS / an issued, unexpired certificate nodes can load (certificates listed as Unusable cannot be chosen; see troubleshooting in HTTPS)No TLSThe certificate of the handshake
Minimum TLS versionTLS 1.2 / TLS 1.3TLS 1.2Handshakes below it are refused
ItemBehavior
SNIMust be one of the certificate's names (a wildcard covers one label), else the handshake is aborted; a handshake without SNI gets the certificate
Cipher suitesThe sites' Modern profile, not configurable
With the PROXY protocolBoth work together: with Accept PROXY protocol the header comes before the handshake; Send to origins still sends one
IP lists and limitsChecked after the handshake
CertificateA certificate an app uses cannot be deleted ("The certificate is still used by L4 applications: …"); after a renewal nodes use the new one without a reload
UDPNot supported ("TLS termination is for TCP applications only")

Verify:

echo NAME | openssl s_client -quiet -connect <node IP>:9000 -servername game.example.com

PROXY protocol

TCP apps only; UDP apps show "UDP does not carry PROXY protocol".

FieldValuesDefaultEffect
Send to originsNone / v1 / v2NoneEvery connection to an origin starts with a PROXY protocol header: the client's address and port, and the node address and port it connected to. v2 headers carry no TLVs
Accept PROXY protocolOn / offOffThe listener expects every connection to start with a PROXY protocol header, v1 or v2 detected automatically; for nodes behind a layer-4 load balancer. The client address is the one in the header (the TCP peer for UNKNOWN), and IP lists check that address
AcceptSend to originsForwarding
OffAnyNative nginx forwarding
OnNoneNative nginx forwarding
Onv1 / v2Lua relay: the header it writes carries the client address from the received header (a header nginx sends natively would carry the load balancer's address). Lower throughput than native forwarding; no half-close, one side closing ends the connection

With Accept PROXY protocol on, connections without a header are closed, and a client that reaches the port directly can claim any address in the header: open the port to the load balancer only. Both settings are structural like the port and protocol: changing them reloads the nodes, see Reloads and long connections.

IP lists and connection limits

FieldValuesDefaultEffect
IP lists · Allow listsAny IP lists, at most 16NoneWhen set, only addresses in these lists are accepted
IP lists · Block listsAny IP lists, at most 16NoneAddresses in these lists are refused
Limits per node · Concurrent connections0–100000000Connections (UDP: sessions) a node forwards at once; 0 means no limit
Limits per node · New connections per second0–10000000New connections (UDP: new sessions) a node accepts per second; 0 means no limit

Add list picks from IP lists; a list's Action does not matter here.

ItemBehavior
OrderAllow lists, block lists, new connections per second, concurrent connections; the first one not met refuses
RefusalNothing is forwarded: TCP connections are closed (data the client sent and nobody read makes the kernel answer RST), UDP datagrams are dropped; counted as Refused. A refused UDP client counts once per datagram
CountingLimits count per node; the cluster's total capacity is the number of nodes × the limit
List changesNodes apply changes to list entries without a reload; a list referenced by an L4 app cannot be deleted ("The IP list is used by …", naming the users)
Other protectionGlobal Block and Allow lists, site bans and HTTP-layer bans, rules, WAF, challenges, and CC protection do not apply to L4 apps. On nodes with kernel-ban-v1, global bans drop packets in the kernel with nftables, L4 ports included, see Kernel bans

DNS

L4 apps use the cluster's DNS steering like sites, see DNS steering and alerts.

ItemBehavior
RecordsOne <app UUID>.<cluster domain> CNAME all.<cluster domain> per enabled app; with Keep per-site line targets also <line name>.<app UUID>.<cluster domain>
SteeringResolution lines, backup node groups, health removal, and scheduling rules apply as usual
HealthNodes are removed by node health (heartbeat, data plane, current revision applied, scheduling rules), not by the origin state of a single app
WritingAfter creating or enabling, the next reconciliation (every minute) writes the record; manual mode lists it among the records to create by hand
DisabledA disabled app has no record; the next reconciliation deletes it (disabled sites keep theirs)
DNS not managedCNAME target shows "Cluster DNS is off"

Clients connect to the CNAME target with the app's port, for example on your own domain:

game.example.com.  CNAME  <app UUID>.<cluster domain>.

Clients connect to game.example.com:9000. CNAME target on the app page also lists the target of every line, for clients that should use one line only.

Statistics

The Analytics tab of the app page (/l4/<app ID>?tab=stats), with the ranges Last hour, Last 6 hours, Last 24 hours (default), and Last 7 days; Refresh reads again.

MetricMeaning
ConnectionsTCP connections or UDP sessions accepted
RefusedConnections or sessions refused by the IP lists or the limits per node
Peak concurrencyPer minute the sum of the nodes' peaks, the highest minute of the range
ReceivedBytes received from clients
SentBytes sent to clients

Charts Connections (connections and refused) and Traffic (received and sent); the Nodes table lists every reporting node's values over the range, busiest first.

ItemBehavior
ReportingNodes report per app and per minute, in the same batch as site statistics; bytes of open connections are sampled every 10 seconds, so long connections count in every minute
RetentionMinute data for 7 days; deleted with the app
LossDeleting a cluster's last L4 app, or an nginx restart, loses the last minute or two a node has not reported yet
Node metricsA node's active connections (metrics-v1) include layer-4 client connections, upstream connections, and UDP sessions, and so does the Active connections condition of scheduling rules

Reloads and long connections

ChangeNode behavior
Port, last port, protocol, whether TLS is terminated, Accept PROXY protocol, the version under Send to origins; creating, deleting, disabling, enablingStructural: nginx.conf is rendered again and nginx reloads
Origins, weights, backups, the origin port mode, passive health check, timeouts, IP lists and their entries, limits per node, the chosen certificate and minimum TLS versionHot update without a reload; new settings apply to later connections
ItemBehavior
TCP connectionsConnections opened before a reload stay with the old worker and keep forwarding until they end or reach the idle timeout; the old worker exits after its last connection
UDP sessionsDo not survive a reload: later datagrams start a new session in a new worker (possibly with another origin); the old session ends at the idle timeout
New portsConnections that arrive between the reload and the push of the L4 app table (milliseconds) are closed; the node reports the configuration applied only after the push
Old workersBy default they serve their connections until these end, so frequent structural changes leave several sets of old workers. The node option --stream-shutdown-timeout makes old workers close the connections they still serve after that long, HTTP keep-alive and WebSocket connections included, see Add nodes

UDP

ItemBehavior
SessionsOne client address and port is one session until the idle timeout; any number of replies per datagram from the origin is forwarded (DNS, QUIC, game protocols work)
Idle timeout30 seconds by default. Set a short idle timeout for request-response protocols such as DNS, otherwise every client port holds a concurrent session until it times out
PROXY protocolNot supported
RefusalRefused datagrams are dropped and counted one by one

Node requirements

ItemBehavior
Capabilityl4-v1 once the configuration has an enabled L4 app
Nodes without itRefuse configurations with L4 apps, keep their last-known-good configuration, and show Upgrade required in Clusters & nodes; since they do not apply the current revision, they leave DNS steering after 2 minutes. The port pools tab, the app list, and the dialog warn "Nodes {nodes} of {cluster} lack L4 forwarding and refuse configurations with L4 apps until upgraded"
OrderUpgrade the nodes first, then create L4 apps
PortsOpen the port pools in the node hosts' firewall and the cloud security group, and publish them on container nodes, see Add nodes
PrivilegesPorts are 1024 or higher: the node's systemd unit and image need no extra privileges

Rollback

Roll back in the cluster's Revisions ships the L4 app settings of the chosen revision:

CaseBehavior
The app is disabled nowNot shipped
The app was deleted since, its port is no longer inside a port pool, or a referenced IP list was deletedThe rollback is refused ("Rollback references resources that are no longer assigned or available")
The configuration canary rolls back to the stable revisionThese apps are left out, everything else is shipped

A rollback does not change the apps' saved settings; any later change publishes them again from the saved settings.

Limits

ItemLimit
AppsAt most 256 per cluster ("A cluster can have at most 256 L4 applications"), disabled ones included
Port rangesAt most 1000 ports per app; 2048 ports per cluster
TLS terminationTCP only; one certificate per app
Ports1024–65535, inside a port pool of the cluster for the protocol; a port of a protocol belongs to one app of the cluster (disabled apps included)
Port poolsAt most 64 per cluster
Origins1–32 per app, at least one not a backup
IP listsAt most 16 allow lists and 16 block lists
PROXY protocolTCP only; v2 headers carry no TLVs; accepting and sending together uses the Lua relay
Passive health stateOn the node only
StatisticsPer minute, kept for 7 days
ProtectionOnly the app's own IP lists, the limits per node, and kernel bans
APIService accounts cannot call the port pool and L4 app procedures, see API and endpoints

Troubleshooting

SymptomCauseFix
"Port {port} is outside the cluster's port pools for this protocol"The port is in no pool of the protocolAdd a port pool on the cluster's Port pools tab, or use another port
"Ports in use by {apps}"Creating or editing: another app of the cluster (disabled ones included) uses the port and protocol; saving port pools: an app's port would be left outsideUse another port, or change or delete the listed apps first
"Port pools overlap: {pools}"Pools of one protocol overlap; TCP + UDP overlaps TCP and UDPChange the ranges or protocols
"A cluster can have at most 256 L4 applications"The limit is reachedDelete unused apps
"PROXY protocol is for TCP applications only"The API set PROXY protocol on a UDP appRemove the PROXY protocol settings
"Nodes {nodes} of {cluster} lack L4 forwarding …"Nodes lack l4-v1Upgrade the nodes
Some nodes leave DNS after an app is createdAs above: these nodes do not apply the current revisionUpgrade the nodes, or disable the app first
Connections time outThe node firewall or security group does not open the port; a container node does not publish itOpen it as in Ports and firewall
Connections are closed at onceRefused by an IP list or a limit (Refused rises); Accept PROXY protocol is on but the client sends no header; the port was just addedCheck the lists and limits; turn on Accept PROXY protocol only behind a load balancer
The origin drops the connection or reports a protocol error right after connectingSend to origins is set but the origin does not accept PROXY protocol, or expects another versionEnable that version on the origin, or select None
The origin sees the load balancer as the clientA layer-4 load balancer sits in front of the nodesHave the load balancer send PROXY protocol, turn on Accept PROXY protocol, and choose a version under Send to origins
Every connection fails and Refused does not riseNo origin is reachableCheck down of GET /v1/l4 on a node, the origin addresses, and the origin's firewall
High UDP Peak concurrency, or the Concurrent connections limit reachedEvery client port is a session until the idle timeoutLower Idle (seconds)
Old nginx workers stay after a reloadThey still serve long connectionsExpected; set --stream-shutdown-timeout to bound it
CNAME target shows "Cluster DNS is off"The cluster's DNS mode is Not managedBind the cluster on its DNS tab
"The IP list is used by …"Deleting a list the listed apps still referenceRemove it from the app's allow and block lists first
Edit on GitHub

On this page