Set Up External Relay Servers

Updated

This guide is part of the Splitting Your Self-Hosted Deployment guide. It covers deploying external relay and STUN servers and configuring your main server to use them.

For each relay server you want to deploy:

Server Requirements

  • A Linux VM with at least 1 CPU and 1 GB RAM
  • Public IP address
  • A domain name pointing to the server (e.g., relay-us.example.com)
  • Docker installed
  • Firewall ports open: 443/tcp and 443/udp (relay), and 3478/udp (STUN). If you configure multiple STUN ports, open all of them

Generate Authentication Secret

All relay servers must share the same authentication secret with your main server. You can generate one with:

# Generate a secure random secret
openssl rand -base64 32

Save this secret. The identical value goes in two settings: NB_AUTH_SECRET on every relay server, and relays.secret on your main server.

Paste it from one source into each file rather than retyping it. A single wrong character produces a relay that looks healthy and refuses every peer.

Create Relay Configuration

On your relay server, create a directory and configuration:

mkdir -p ~/netbird-relay
cd ~/netbird-relay

Create relay.env with your relay settings. The relay server can automatically obtain and renew TLS certificates via Let's Encrypt:

NB_LOG_LEVEL=info
NB_LISTEN_ADDRESS=:443
NB_EXPOSED_ADDRESS=rels://relay-us.example.com:443
NB_AUTH_SECRET=your-shared-secret-here

# TLS via Let's Encrypt (automatic certificate provisioning)
NB_LETSENCRYPT_DOMAINS=relay-us.example.com
NB_LETSENCRYPT_EMAIL=admin@example.com
NB_LETSENCRYPT_DATA_DIR=/data/letsencrypt

# Embedded STUN (comma-separated for multiple ports, e.g., 3478,3479)
NB_ENABLE_STUN=true
NB_STUN_PORTS=3478

The file holds the shared secret, so make it readable by root only:

chmod 600 relay.env

The relay reads relay.env when its container is created. If you change the file later on a running relay, apply it with docker compose up -d. docker compose restart keeps the old values.

Create docker-compose.yml:

services:
  relay:
    image: netbirdio/relay:latest
    container_name: netbird-relay
    restart: unless-stopped
    ports:
      # Both relay transports: WebSocket over TCP and QUIC over UDP.
      # Docker treats a bare '443:443' as TCP only, so the UDP line is required.
      - '443:443/tcp'
      - '443:443/udp'
      # Expose all ports listed in NB_STUN_PORTS
      - '3478:3478/udp'
    env_file:
      - relay.env
    volumes:
      - relay_data:/data
    logging:
      driver: "json-file"
      options:
        max-size: "500m"
        max-file: "2"

volumes:
  relay_data:

Running the Relay Behind a Proxy or Load Balancer

In this shape your proxy holds the certificate and the relay serves plain HTTP behind it.

Use this shape only because a load balancer or proxy is already in front of your relays. Do not add one just to terminate TLS on the relay host: the relay already does that itself, and a proxy in front of it costs you QUIC.

On the relay host, relay.env has no NB_LETSENCRYPT_* lines, listens on an internal port, and names the proxy's address:

NB_LOG_LEVEL=info
NB_LISTEN_ADDRESS=:8080
NB_EXPOSED_ADDRESS=rels://relay-us.example.com:443
NB_AUTH_SECRET=your-shared-secret-here
NB_TRUSTED_PROXIES=10.20.0.10
NB_ENABLE_STUN=true
NB_STUN_PORTS=3478

Keep rels:// written out in NB_EXPOSED_ADDRESS. With no TLS of its own, the relay would otherwise advertise rel://.

In docker-compose.yml, remove the two 443 port lines and keep 3478:3478/udp for STUN. Make port 8080 reachable from the proxy and nothing else. If the proxy is on another machine, publish the port with '8080:8080/tcp' and allow it from the proxy's address only, at a network firewall or cloud security group. A host firewall such as UFW or firewalld does not restrict it, because Docker-published ports bypass those rules.

Your proxy must:

  • hold a certificate for the relay's domain, and accept connections on 443/tcp;
  • forward WebSocket upgrades over HTTP/1.1 to the relay's port 8080;
  • set both X-Real-Ip and X-Real-Port to the address and source port of the connection it received, overwriting anything the client sent.

The relay uses those two headers only when the connection comes from an address in NB_TRUSTED_PROXIES, and only when both are present. A proxy that sets only X-Real-Ip gets its own address logged for every peer. A proxy that passes the client's headers through lets any client write the address of its choice into your relay logs.

NB_TRUSTED_PROXIES is the address the relay sees the proxy's connections come from, which is not always the proxy's public address. When the proxy and relay share a Docker network, it is the proxy container's address on that network. When the proxy is another machine on the same private network, it is that machine's private address: a published Docker port keeps the original source address.

These headers decide only which client address the relay records, including in the invalid signature lines used for troubleshooting. Authentication does not depend on them.

Apply a change to NB_TRUSTED_PROXIES with docker compose up -d relay. An invalid entry stops the relay: it exits with failed to parse trusted proxies: ... and restarts in a loop until you fix it. To confirm the setting works, check that the relay logs your peers' addresses rather than the proxy's:

docker compose logs relay | grep 'WS client connected from'

The CLI equivalent of NB_TRUSTED_PROXIES is --trusted-proxies, available since v0.75.0.

At startup the relay reports that it cannot serve QUIC:

WARN relay/server/server.go:78: Not starting QUIC listener: valid TLS config is required for QUIC listener

That is expected here. Peers connect over WebSocket, so neither the proxy nor the relay host needs 443/udp. The proxy needs 443/tcp, plus whatever it uses to obtain its own certificate, often 80/tcp. The relay host needs 3478/udp for STUN.

Alternative: TLS with Existing Certificates

If you have existing TLS certificates (e.g., a wildcard certificate, or one from your own CA), delete the three NB_LETSENCRYPT_* lines from relay.env and add:

NB_TLS_CERT_FILE=/certs/fullchain.pem
NB_TLS_KEY_FILE=/certs/privkey.pem

Then replace the relay service's volumes: list in docker-compose.yml with this one. Do not add it as a second volumes: key, which makes the file invalid:

    volumes:
      - /path/to/certs:/certs:ro
      - relay_data:/data

fullchain.pem is your certificate followed by any intermediates, and the certificate must cover the relay's domain. Four things catch people out:

  • Set both variables, and remove every NB_LETSENCRYPT_* line. If a Let's Encrypt line remains, the relay uses Let's Encrypt and ignores your files, and the curl -v check below shows a Let's Encrypt issuer instead of yours. If only one NB_TLS_* variable is set, the relay starts with no TLS at all, and the sign is a missing QUIC line at startup.
  • Nothing in the startup log names your certificate. The curl -v check below is how you confirm it is the one being served.
  • The files are read once, at startup. Restart the relay after every renewal: docker compose restart relay.
  • A private CA must be trusted by every peer. NetBird checks the relay's certificate against each device's operating-system trust store, and has no setting for a separate CA. A device that does not trust yours reports the relay as Unavailable with x509: certificate signed by unknown authority, and cannot relay to other peers. Distribute the CA to every device before you switch.

Start the Relay Server

docker compose up -d

Verify it's running:

docker compose logs -f

You should see the relay announce its exposed address, both listeners, and the STUN server:

INFO relay/cmd/root.go:242: server will be available on: rels://relay-us.example.com:443
INFO relay/server/listener/ws/listener.go:51: WS server listening address: :443
INFO relay/server/listener/quic/listener.go:39: QUIC server listening on address: :443
INFO [component: stun] stun/server.go:71: STUN server listening on [::]:3478

Other lines appear alongside these, including the Let's Encrypt setup, the health check server and the metrics server. The order changes from one start to the next, because the listeners come up concurrently. What matters is that all four lines are present, not where they sit.

A missing QUIC line means the relay has no TLS configuration of its own, which is expected only behind a TLS-terminating proxy.

If you configured Let's Encrypt, the relay generates TLS certificates lazily on the first incoming request. Trigger certificate provisioning and verify it by running:

curl -v https://relay-us.example.com/

A 404 page not found response is expected. What matters is that the TLS handshake succeeds:

* Server certificate:
*  subject: CN=relay-us.example.com
*  issuer: C=US; O=Let's Encrypt; CN=E8
*  SSL certificate verify ok.

The issuer's CN names whichever intermediate signed your certificate, so yours will often differ.

Two things can go wrong here, and they look different. If the first attempt fails with a TLS error such as SSL_ERROR_SYSCALL, wait a few seconds and run it again: that first request is what triggers issuance, and it can time out while that happens.

If instead the command hangs with no output and no error, issuance is stuck rather than slow, and only the relay can tell you why:

docker compose logs relay | grep -i acme

A rate limit is the likeliest cause if you have rebuilt the same relay host several times. Let's Encrypt allows five certificates per week for one exact set of domain names.

Repeat for Additional Relay Servers

If deploying multiple relays (e.g., for different regions), repeat the steps above on each server. Use the same NB_AUTH_SECRET but update the domain name for each.

Update Main Server Configuration

Now update your main NetBird server to use the external relays instead of the embedded one.

Edit config.yaml

On your main server, edit the config.yaml file:

cd ~/netbird  # or wherever your deployment is
nano config.yaml

Add stuns and relays inside the existing server: block, at the same indentation as the keys already there. Anywhere inside that block works.

This is an addition, not a replacement. Leave every key already under server: exactly as it is. A quickstart deployment also carries auth, reverseProxy and store sections there, holding your OIDC redirect URIs, your reverse-proxy trust settings and your datastore encryption key. Overwriting any of them breaks the deployment, so do not paste the example below over your existing block.

The presence of relays.addresses is what disables the embedded relay, and it disables the embedded STUN server too, so the stuns section is required to provide external STUN addresses.

You do not need to remove server.authSecret. It is required only when the embedded relay is running. Leaving it in place is also the safer choice on a NetBird Enterprise deployment, where the same value has a second job: see External Relays on a Commercial License.

Both relays and stuns are available since v0.65.0.

server:
  listenAddress: ":80"
  exposedAddress: "https://netbird.example.com:443"
  # Leave authSecret as it is. relays.addresses below is what disables the embedded relay.
  authSecret: "your-existing-secret"
  # stunPorts no longer applies once stuns is set, so you can comment it out
  # stunPorts:
  #   - 3478
  metricsPort: 9090
  healthcheckAddress: ":9000"
  logLevel: "info"
  logFile: "console"
  dataDir: "/var/lib/netbird"

  # External STUN servers (your relay servers)
  stuns:
    - uri: "stun:relay-us.example.com:3478"
      proto: "udp"
    - uri: "stun:relay-eu.example.com:3478"
      proto: "udp"

  # External relay servers
  relays:
    addresses:
      - "rels://relay-us.example.com:443"
      - "rels://relay-eu.example.com:443"
    secret: "your-shared-secret-here"
    credentialsTTL: "24h"

  # ... the rest of your existing configuration is unchanged

A mismatch is easiest to spot on the relay host, which logs failed to handshake: ... invalid signature for every rejected peer.

It is harder to spot from a peer, because a peer lists only the relays it is actually using. One bad relay among several never appears as a failure: the peer picks a working one instead, and the broken relay is simply absent. Only when every relay rejects does netbird status -d report Unavailable, reason: failed to get reader: failed to read frame header: EOF. The stun: entry stays Available throughout, because STUN is unauthenticated.

Update docker-compose.yml (Optional)

If your main server was exposing STUN port 3478, you can remove it, since STUN is now handled by external relays. Change only the ports entry and leave the rest of the service as it is:

  netbird-server:
    image: netbirdio/netbird-server:latest
    container_name: netbird-server
    restart: unless-stopped
    networks: [netbird]
    # Remove the STUN port - no longer needed
    # ports:
    #   - '3478:3478/udp'
    volumes:
      - netbird_data:/var/lib/netbird
      - ./config.yaml:/etc/netbird/config.yaml
    command: ["--config", "/etc/netbird/config.yaml"]
    # ... the rest of this service, including its labels, is unchanged

Restart the Main Server

Which command you need depends on whether you did the optional docker-compose.yml step above.

If you edited config.yaml only, restart the container:

docker compose restart netbird-server

config.yaml is a bind mount, so editing it does not change anything Docker Compose compares. docker compose up -d would report the service as already up to date, and the server would keep serving the old configuration.

If you also removed the STUN port from docker-compose.yml, recreate the container instead:

docker compose up -d netbird-server

A published port is fixed when the container is created, so docker compose restart cannot remove one: the container comes back with the mapping still in place and nothing reports a problem. Recreating the container applies both files at once, so this one command covers the config.yaml change as well.

The Embedded Relay Is Now Off

Setting relays.addresses switches the embedded relay off. There is no setting that keeps it running alongside external relays, and listing the main server's own address among them does not bring it back: clients are handed an address that never answers.

curl https://netbird.example.com/relay
Relay service not enabled

The endpoint answers with HTTP 404 and that single line as its body. Add -i to the command if you want to see the status code as well.

From here on, relay capacity is whatever your external relay hosts provide. Size them accordingly.

Verify the Configuration

Check Main Server Logs

docker compose logs netbird-server

Verify that the embedded relay is disabled and your external relay addresses are listed:

INFO combined/cmd/root.go:   Management: true (log level: info)
INFO combined/cmd/root.go:   Signal: true (log level: info)
INFO combined/cmd/root.go:   Relay: false (log level: )
Relay addresses: [rels://relay-us.example.com:443 rels://relay-eu.example.com:443]

The address list is logged twice, from two different places in the server. That is normal.

Check Peer Status

Connect a NetBird client and verify that both STUN and relay services are available:

netbird status -d

The output lists your external STUN and relay servers. Every configured STUN server appears. Relays work differently: the client dials them in parallel and keeps the first to answer, which is normally the nearest. That one becomes its home relay, so relays in several regions give each client the nearest one automatically.

Relays:
  [stun:relay-us.example.com:3478] is Available
  [stun:relay-eu.example.com:3478] is Available
  [rels://relay-eu.example.com:443] is Available via ws

You may see more than one rels:// line. A client also connects to another peer's home relay when that peer picked a different one, so holding two relay connections at once is normal in a multi-region deployment.

The suffix names the transport that won the race, and it is ws or quic depending on which answered first. Both are normal, the winner can differ between two peers on the same deployment, and neither is a sign of a problem. What 443/udp buys you is that the QUIC attempt can win or lose on merit rather than sitting until it times out.

Test Failover

Stop the relay a peer is actually using, which is the one on its rels:// line, otherwise nothing observable changes. Then check that peer again.

Read the rels:// line, not the stun: lines. It is the one that shows failover working: the stopped relay drops out of it and the surviving relay carries the traffic.

  [rels://relay-us.example.com:443] is Available via ws

The stun: entries are not a dependable failover signal. Depending on circumstances they either report the stopped server as Unavailable with a reason, or sit at Checking... for every configured STUN server, including healthy ones, without clearing on their own or after a client restart. Both have been observed. Either way they return to Available once the stopped relay is back, so do not read Checking... as a second failure.

Test Relay Connectivity

You can force all peer connections through relay to verify it works end-to-end. On a client, run:

sudo netbird service reconfigure --service-env NB_FORCE_RELAY=true

Then test connectivity to another peer (e.g., with ping).

Once confirmed, switch back to normal mode. The client will attempt peer-to-peer connections first and fall back to relay only when direct connectivity isn't possible:

sudo netbird service reconfigure --service-env NB_FORCE_RELAY=false