Running a Highly Available Self-Hosted Deployment

Updated

A single NetBird server works well for many deployments, but production environments may need to remain available when an individual host or service instance fails. A highly available deployment removes those single points of failure by spreading services across separate failure domains and routing traffic to healthy instances.

NetBird Enterprise supports active-active high availability (HA) for Management and Signal. This guide shows how to move an existing self-hosted deployment to that mode: Management, Signal, and Relay run as independent pools: Management and Signal behind load balancers, and Relay with an address of its own for every instance. PostgreSQL, Redis, and NATS provide the shared state that lets a healthy Management or Signal instance continue serving traffic when another instance fails.

Use this guide when you need to tolerate the loss of a Management, Signal, Relay, or NATS node and perform rolling upgrades with minimal disruption. It assumes an existing Enterprise deployment that already uses PostgreSQL. For a distributed deployment without active-active Management and Signal, see Splitting Your Self-Hosted Deployment.

Before you begin

This is an infrastructure migration, not an in-place switch. Build and validate the new dependencies before changing the public endpoints used by your existing deployment.

ConfirmWhy it matters
You have a recent, restorable PostgreSQL backup.Database migration or endpoint changes should always have a rollback path.
You have recorded the current server.store.encryptionKey value.Every Management replica must use the exact same key to read existing encrypted data.
You can deploy instances in separate failure domains.Two containers on one host do not provide high availability.
You have stable DNS names and load balancers for Management and Signal, and a DNS name for every Relay instance.Peers and the dashboard must keep using stable URLs while backends change. Relay instances are reached by their own names, not through a load balancer.
PostgreSQL and Redis can each expose one highly available, read-write endpoint.Management replicas need a shared database and cache, not per-instance stores.
You can test the change in staging first.Failure testing is part of validating an HA deployment.

Deployment order

Build the HA deployment in this sequence:

  1. Make PostgreSQL and Redis highly available.
  2. Deploy and validate the three-node NATS cluster.
  3. Create the Management and Signal load-balancer frontends, and the DNS records for every service, including one per Relay instance.
  4. Deploy the Relay and Signal pools.
  5. Configure and deploy the Management pool.
  6. Validate normal operation and then run the failure tests.

Scope and endpoint choices

This guide uses these public service URLs:

ServiceExample URLUsed by
Managementhttps://netbird.example.comDashboard, Management API, Management gRPC, and OAuth
Signalhttps://signal.example.comNetBird peers for signaling
RelayOne per instance, e.g. rels://us-1.relay.example.com:443 and rels://eu-1.relay.example.com:443. With the geo-DNS option, also a shared rels://relay.example.com:443.NetBird peers that need relay connectivity

The Management URL can remain the URL used by your existing deployment. If you change it, update the dashboard environment to use the new Management API and gRPC endpoints; see Dashboard environment configuration. The dashboard itself is stateless and can be served behind the Management load balancer or separately, as long as browsers can reach the configured Management URL.

Architecture

A highly available deployment has three service pools. They have different connection patterns, so they scale and fail over independently instead of forcing every host to handle every workload.

  • Management pool: Enterprise Management replicas serve the dashboard API, Management gRPC, and OAuth2 endpoints. Fully stateless: any replica can serve any request because all durable state lives in PostgreSQL and all short-lived state lives in Redis. Replicas scale horizontally to absorb dashboard traffic, peer sync, and OAuth flows.
  • Signal pool: Enterprise Signal replicas serve the Signal gRPC API. Peers connect to one Signal instance through the load balancer. The NATS cluster reconciles cross-instance peer signaling, so any Signal instance can deliver a message to any peer regardless of its connected instance. This makes Signal active-active in the Enterprise build. See How active-active Signal works.
  • Relay pool: Relay instances carry traffic for peers that cannot reach each other directly. Unlike Management and Signal, Relay does not go behind a load balancer. Relay instances share no state, so two peers can be relayed to each other only when each can reach the other's instance by that instance's own address. Every instance therefore has its own URL, and a peer that loses its relay moves to another instance by itself. See Step 5.

If you use traffic events, two more services run alongside the Management pool: a flow receiver and a flow enricher. See Deploy the flow receiver and enricher.

The pools depend on the following shared infrastructure:

  • NATS cluster: Routes cross-instance Signal messages, distributes dynamic configuration (log level and rate limits), and carries the traffic-flow event stream. Quorum is mandatory. Loss of quorum stalls cross-instance signaling.
  • PostgreSQL HA endpoint: Stores durable Management data, including accounts, peers, policies, OAuth state, and integrations. It is operator-managed. NetBird treats it as an opaque connection string and expects failover to happen transparently behind the DSN.
  • Redis HA endpoint: Stores ephemeral cache data, including OAuth PKCE verifiers, peer cache data, and dynamic configuration. It is operator-managed. Losing Redis interrupts in-flight OAuth flows but does not break running peer connections.

The Management and Signal pools each use one stable URL backed by load-balanced instances, so adding or removing one of their instances is transparent to peers. The Relay pool is reached by each instance's own address instead.

The two load balancers can be independent load balancers, frontends on one shared load balancer, or managed load-balancer resources. Choose the model that fits your infrastructure. Management and Signal must each be reached through one stable URL with at least two backend instances and a load balancer that can fail over within seconds.

How active-active Signal works

In single-node mode, the Signal service keeps peer connection state in memory and cannot be replicated. In the Enterprise HA build, every Signal instance connects to the NATS cluster and uses it to route signaling messages between peers connected to different instances. When peer A, connected to Signal 1 through the load balancer, needs to reach peer B, connected to Signal 2, Signal 1 publishes the signaling message to NATS. Signal 2's NATS subscription forwards it to peer B. The Signal load balancer can distribute peers across instances freely because NATS reconciles cross-instance routing. A healthy NATS cluster is mandatory.

Prerequisites

  • An active NetBird Enterprise commercial license.
  • At least 2 enterprise Management instances, on separate failure domains.
  • At least 2 enterprise Signal instances, on separate failure domains.
  • At least 2 Relay instances, on separate failure domains.
  • At least 3 NATS instances for the coordination cluster, on separate failure domains. NATS can colocate with NetBird hosts, but the 3 NATS instances must be on different failure domains.
  • A load balancer for the Management and Signal pools. These can be two independent load balancers, two frontends on one shared load balancer, or two managed load-balancer resources, as long as each pool is reachable through a single stable URL. Both require HTTP/2 + gRPC support. The Relay pool does not use a load balancer.
  • Public FQDNs for Management and Signal, for example netbird.example.com and signal.example.com, each resolving to its load balancer. One more per Relay instance, for example us-1.relay.example.com and eu-1.relay.example.com, each resolving to that instance. With the geo-DNS option, also a shared relay.example.com.
  • Permissions to deploy services, mount configuration and secrets, expose network ports, manage DNS records, and register instances with the load balancers in your environment.

Step 1: Make PostgreSQL highly available

NetBird treats PostgreSQL as an opaque connection string. You give it a DSN; NetBird makes no assumptions about replication, failover, or topology behind that DSN. PostgreSQL HA is your responsibility.

What NetBird requires

  • A single, stable connection endpoint that survives node failure. NetBird does not implement read replicas, partitioning, or client-side connection pooling against a list of hosts.
  • Read-write access on every store DSN. NetBird writes on every store; a read-only replica is not sufficient.
  • All three NetBird stores (server.store, server.activityStore, server.authStore) configured with PostgreSQL DSNs. They can point at the same database/instance or at separate ones, depending on your isolation requirements.
  • TLS is recommended (sslmode=require in the DSN).

Common HA patterns

PatternNotes
Managed PostgreSQL (AWS RDS Multi-AZ, GCP Cloud SQL HA, Azure Database for PostgreSQL, Aiven, Neon)Easiest. The service exposes a single endpoint and handles failover internally.
Streaming replication + Patroni or Stolon, fronted by PgBouncer or HAProxySelf-hosted. Patroni handles primary election; PgBouncer/HAProxy exposes a single virtual endpoint that follows the elected primary.
PostgreSQL with pg_auto_failoverLighter-weight alternative to Patroni.

For the full PostgreSQL high-availability documentation, see PostgreSQL: High Availability, Load Balancing, and Replication.

If your current PostgreSQL is not HA

If your current PostgreSQL deployment is not highly available, move it behind an HA endpoint before adding multiple Management replicas. Use the database platform your organization already trusts, such as a managed PostgreSQL service, a Patroni or Stolon cluster, pg_auto_failover, or another standard HA pattern. NetBird only depends on the endpoint behavior, so the implementation can follow your existing database operations model. Whichever approach you choose, the constraints below apply.

At a minimum, your migration must:

  1. Take a logically consistent backup of the current PostgreSQL data. For example, use pg_dump, a snapshot, or a point-in-time copy.
  2. Restore the data into the new HA PostgreSQL setup, preserving schemas and table ownership.
  3. Update server.store.dsn, server.activityStore.dsn, and server.authStore.dsn in config.yaml to point at the new HA endpoint.
  4. Keep server.store.encryptionKey identical to the value used by the original deployment.

Validate after migration by starting one replica against the new endpoint and confirming it boots cleanly, the dashboard loads, and existing peers reconnect.

How Management points at PostgreSQL

Every Management replica's config.yaml references the PostgreSQL HA endpoint through three store DSNs. All three can use the same database (or separate databases if you prefer isolation between management, activity, and auth stores):

server:
  store:
    engine: "postgres"
    dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"
    encryptionKey: "<preserved from existing deployment; identical on every replica>"

  activityStore:
    engine: "postgres"
    dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"

  authStore:
    engine: "postgres"
    dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"

The full config.yaml is assembled in Step 7.

Step 2: Set up a Redis HA endpoint

NetBird uses Redis as a shared cache across the Management pool. Redis does not hold persistent NetBird state. If Redis becomes unavailable, in-flight OAuth flows fail until it returns. Established peer connections continue working because they do not query Redis for every message.

What NetBird requires

NetBird connects to Redis as a single standalone instance using one URL and the standard Redis protocol. The URL must point to a stable endpoint that handles failover internally. NetBird does not negotiate failover itself.

  • URL schemes: redis:// or rediss:// (TLS). TLS is strongly recommended for any network-traversing traffic.
  • One URL only. NetBird does not accept a list of hosts and does not perform client-side failover. Whatever sits behind the URL is responsible for failover.
  • No cluster-aware or Sentinel-aware behavior. NetBird does not negotiate the Redis Cluster protocol (MOVED redirects, slot-aware routing) or the Redis Sentinel discovery protocol. The endpoint must look like a standard standalone Redis instance from NetBird's side.

What works for HA

ApproachWorks?Notes
Managed Redis with a single primary endpoint: AWS ElastiCache for Redis (cluster mode disabled / replication group with replicas), GCP Memorystore for Redis (Standard tier), Azure Cache for Redis (Standard or Premium tier without clustering), Upstash Redis (regional), Redis Enterprise Cloud (active-passive)YesEasiest path. The service exposes one DNS name and handles failover internally. The DNS name follows the new primary automatically.
Self-hosted Redis Sentinel + TCP proxy (HAProxy with a Sentinel-aware backend selector, Nginx stream module, or a keepalived-managed VIP that follows the Sentinel-elected master)YesSentinel elects the primary. The TCP proxy exposes one stable address that forwards to the current primary. NetBird connects to the proxy and does not communicate with Sentinel.
Redis Cluster (cluster mode enabled, multi-shard) even with a single-VIP front doorNoNetBird keys can hash to different cluster slots. Only keys served by the shard behind the VIP are reachable. Other reads and writes fail with MOVED errors that NetBird cannot follow.
Multiple Redis URLs (round-robin DNS, comma-separated list, etc.)NoNetBird takes a single URL string. Failover must be transparent below the URL, not negotiated client-side.
  • If you run on a major cloud: Use managed Redis in non-cluster / standalone-with-replication mode. Look for "cluster mode disabled", "replication group", or "standard tier with failover" in your provider's UI. This option has the least operational burden because the provider hides failover behind a single DNS name.
  • If you self-host: Redis Sentinel + TCP proxy is the right pattern. See the Redis Sentinel documentation for Sentinel setup. For the proxy layer, HAProxy with a Sentinel-aware health check (or keepalived for an IP-level VIP) is the typical choice.

What gets stored

DataTTL / pattern
OAuth PKCE verifiers, one-time proxy tokens10 minutes
Peer cache (public-key → account/peer ID)30 minutes + jitter
Dynamic log-level and rate-limit configurationPolled every 5 minutes, pub/sub for instant updates

If Redis becomes unreachable, in-flight OAuth flows fail. Established peer connections continue working until cache entries expire.

How Management points at Redis

Every Management replica's config.yaml references the Redis HA endpoint through server.ha.redisAddr:

server:
  ha:
    enabled: true
    redisAddr: "redis://redis.example.com:6379"
    # natsEndpoints set in Step 3 below

The full config.yaml is assembled in Step 7.

Step 3: Set up a NATS HA cluster

NetBird uses NATS for coordination between Management and Signal instances. Like PostgreSQL and Redis, the NATS deployment is operator-owned: you can use a managed NATS service, a Kubernetes operator, or a self-hosted cluster, as long as it exposes the behavior NetBird requires.

What NetBird requires

  • A highly available NATS cluster with at least three nodes across separate failure domains.
  • JetStream enabled with persistent storage.
  • Quorum available for writes. In a three-node cluster, at least two nodes must be healthy.
  • Client endpoints reachable from every Management and Signal instance.
  • TLS and authentication appropriate for your network boundary.
  • A replicated traffic-events JetStream stream for traffic-flow events.
NATS portPurposeExposure
4222/tcpClient connections from NetBird and Signal instancesReachable from the Management and Signal pools
6222/tcpCluster routes (NATS to NATS)Reachable between NATS nodes only, if you self-host the cluster
8222/tcpMonitoring / health endpointsInternal observability, if enabled by your NATS deployment

Common HA patterns

PatternNotes
Managed NATS with JetStreamLowest operational burden. The provider is responsible for node placement, storage, upgrades, and failover.
NATS Kubernetes operatorGood fit when NetBird runs in Kubernetes or your platform team already operates stateful services there. Use persistent volumes and spread pods across zones or nodes.
Self-hosted NATS clusterWorks well when you already operate VMs or bare-metal services. Place each node in a separate failure domain and back JetStream with persistent storage.

Example: a self-hosted cluster with Docker Compose

This example runs one NATS node per host, on three hosts: nats-1.example.com, nats-2.example.com and nats-3.example.com. A NATS node can share a host with a NetBird node or a Signal instance, but no host may run two NATS nodes.

Generate two passwords, one for NetBird and Signal to connect with and one for the NATS nodes to connect to each other. Use hex, so they are safe inside URLs:

openssl rand -hex 24

On each host, create a directory with two files. The first, nats-server.conf, differs per host in server_name and client_advertise:

server_name: nats-1
port: 4222
client_advertise: "nats-1.example.com:4222"
http_port: 8222

jetstream {
  store_dir: /data
}

authorization {
  user: netbird
  password: "<client password>"
}

cluster {
  name: netbird
  port: 6222
  authorization {
    user: route
    password: "<route password>"
  }
  routes: [
    "nats-route://route:<route password>@nats-1.example.com:6222"
    "nats-route://route:<route password>@nats-2.example.com:6222"
    "nats-route://route:<route password>@nats-3.example.com:6222"
  ]
}

client_advertise is the address each node gives clients for reconnecting. Without it, a node in a container advertises its Docker-internal address, which NetBird and Signal cannot reach.

The second file, docker-compose.yml, is the same on every host:

services:
  nats:
    image: nats:2.14
    container_name: nats
    restart: unless-stopped
    command: ["-c", "/etc/nats/nats-server.conf"]
    ports:
      - "4222:4222"
      - "6222:6222"
      - "127.0.0.1:8222:8222"
    volumes:
      - ./nats-server.conf:/etc/nats/nats-server.conf:ro
      - nats-data:/data

volumes:
  nats-data:

The configuration holds both passwords, so keep it readable by root only, then start the node:

chmod 600 nats-server.conf
docker compose up -d

Allow port 4222 from the Management and Signal hosts only, and port 6222 between the NATS hosts only. Set these rules on your network firewall or security groups: ports that Docker publishes bypass host firewalls such as UFW and firewalld. The monitoring port 8222 listens on the host itself only.

Once all three nodes run, check each one:

curl -s http://127.0.0.1:8222/healthz
curl -s http://127.0.0.1:8222/jsz | grep -E '"cluster_size"|"leader"'

Each node answers {"status":"ok"}, reports "cluster_size": 3 and names the same leader. Until all three have started, the nodes log Error trying to connect to route for the ones that are not up yet.

NetBird and Signal connect as the netbird user, so the NATS URLs you give them carry the client password: nats://netbird:<client password>@nats-1.example.com:4222. Both write these URLs to their logs at startup, password included. Keep those logs as private as the configuration.

This example does not configure TLS, so the passwords and all NATS traffic cross the network unencrypted. Keep NATS on a private network, or add TLS as described in the NATS documentation.

The nats CLI commands below also need the credentials: add --user netbird --password "<client password>". The server report command needs a system account, which this example does not create; use the curl checks above instead.

Traffic-flow stream

NetBird expects the traffic-events JetStream stream to be available for traffic-flow events. Configure it with:

SettingValue
Stream nametraffic-events
Subjectstraffic-events.>
StorageFile-backed persistent storage
Replicas3
RetentionLimits-based
Max age168h
Discard policyOld messages first

If you manage NATS directly, the equivalent nats CLI command is:

nats --server nats://nats-1.example.com:4222 \
  stream add traffic-events \
  --subjects "traffic-events.>" \
  --storage file \
  --replicas 3 \
  --retention limits \
  --max-age 168h \
  --discard old \
  --defaults

The flow receiver publishes to netbird.flow.events by default, a subject this stream does not match. Step 7 sets it to one under traffic-events.; see Deploy the flow receiver and enricher.

If your NATS platform manages streams declaratively, apply the same settings through that platform instead. If the stream already exists with one replica, update it to three replicas:

nats stream edit traffic-events --replicas 3

Verify the cluster

Use your NATS platform's health checks to confirm the cluster is healthy and JetStream has quorum. If you use the nats CLI, a typical check is:

nats --server nats://nats-1.example.com:4222 server report jetstream

You should see three nodes, one elected as JetStream meta-leader, and the traffic-events stream showing three replicas with all peers healthy.

How Management points at NATS

Every Management replica's config.yaml references the NATS cluster through server.ha.natsEndpoints, a comma-separated list of all cluster endpoints. Each Management and Signal instance connects to one of them and handles failover and reconnection across the rest internally:

server:
  ha:
    enabled: true
    natsEndpoints: "nats://nats-1.example.com:4222,nats://nats-2.example.com:4222,nats://nats-3.example.com:4222"
    # redisAddr from Step 2 also goes here

Signal instances reference the same list via the NATS_ENDPOINTS environment variable (see Step 6). The full Management config.yaml is assembled in Step 7.

Why an explicit list and not a single DNS round-robin URL? A single hostname with multiple A/AAAA records (e.g. nats://nats.example.com:4222) works because the underlying client resolves and reconnects, but recovery time depends on DNS TTL and cache. An explicit comma-separated list of node addresses gives immediate visibility of every node and the fastest possible failover when one becomes unreachable. Prefer the explicit list.

Step 4: Configure the load balancers

Use the load balancer your organization already operates. NetBird works with any Layer 7 load balancer that supports HTTP/2 and gRPC, including HAProxy, NGINX, AWS Application Load Balancer, Google Cloud Load Balancing, Azure Application Gateway, and Envoy.

Plan one load-balancer frontend for each of the Management and Signal pools. These can be separate load balancers or two frontends on a shared load balancer. Each frontend needs its own DNS name and backend pool. Configure the frontends and empty backend pools now, then add instances as you deploy the Signal and Management services in Steps 6 and 7. The Relay pool does not go behind a load balancer; see Step 5.

PoolPublic FQDNBackend portFrontend protocolHealth check
Signale.g. signal.example.com443HTTPS, HTTP/2, gRPCTCP/443 (or gRPC health if supported)
Managemente.g. netbird.example.com443HTTPS, HTTP/2, gRPCGET /oauth2/.well-known/openid-configuration → HTTP 200

Common requirements for every pool:

RequirementNotes
TLS terminationPrefer terminating TLS at the load balancer. Set server.tls to empty on the Management replicas. If your environment requires TLS pass-through, configure TLS on every backend instead.
HTTP/2 + gRPC support end-to-endRequired for Management and Signal pools. Make sure your load balancer supports long-lived connections.
Connection-level affinityThis is the default for Layer 4 load balancers and Layer 7 load balancers that use HTTP/2 streaming. A long-lived connection stays on one backend for its lifetime. No application-level sticky sessions are required. NATS reconciles cross-instance Signal state, while PostgreSQL and Redis hold Management state.
Active health checksMark backends out of the pool on failure within seconds, not minutes.
Connection draining on rolling upgradeWhen you remove a backend, the LB should let in-flight gRPC streams finish before tearing them down.
Generous idle timeoutManagement and Signal gRPC streams can be long-lived. Set the idle timeout above your peer-sync interval. Ten to 30 minutes is a comfortable range for both.

For the full set of paths and protocols NetBird exposes to each load balancer, see Configuration Files Reference.

The Management check confirms that a replica is up and serving. It keeps returning 200 while PostgreSQL or Redis is unreachable, so monitor those separately.

If you're using a managed cloud load balancer, configure the equivalent of each row above using your provider's UI or infrastructure as code. If you're using a self-hosted reverse proxy, create one backend pool per service, attach health checks, and enable HTTP/2 or WebSocket support as appropriate for each pool.

Step 5: Deploy the Relay pool

Deploy at least two Relay instances on separate failure domains. The Relay pool uses the upstream Relay image (netbirdio/relay). There is no Enterprise-specific Relay image. Relay does not validate a license. It authenticates incoming connections against the shared NB_AUTH_SECRET.

Relay instances share no state with each other, which is why this pool works differently from Management and Signal. The instance a peer connects to becomes its home relay, and announces its own address to that peer. When two peers are on different instances, each reaches the other's instance directly by that announced address. So every instance needs an NB_EXPOSED_ADDRESS of its own, with a DNS name and a TLS certificate that peers can reach.

Choose how peers find an instance:

Optionserver.relays.addressesWhich instance a peer usesChoose it when
1. List every instanceEvery instance's URLThe first to answer: the client connects to all of them in parallel and keeps the fastest, normally the nearestBy default. There is nothing extra to run.
2. Geo-DNSOne shared name that your DNS provider resolves to a nearby instanceThe one DNS returns for the peer's locationYou run many instances across regions and do not want every peer racing all of them.

In both options, a peer that loses its relay moves to another instance by itself, and peers on different instances still reach each other.

Option 1: list every instance

Give each instance its own DNS name and let it obtain its own Let's Encrypt certificate for that name. On us-1, create /etc/netbird/relay/relay.env:

NB_LOG_LEVEL=info
NB_LISTEN_ADDRESS=:443
# This instance's own address. Different on every instance.
NB_EXPOSED_ADDRESS=rels://us-1.relay.example.com:443
NB_AUTH_SECRET=<shared secret, identical on every Relay instance and on the Management pool>
NB_LETSENCRYPT_DOMAINS=us-1.relay.example.com
NB_LETSENCRYPT_EMAIL=admin@example.com
NB_LETSENCRYPT_DATA_DIR=/data/letsencrypt
NB_ENABLE_STUN=true
NB_STUN_PORTS=3478

The file holds the shared secret, so make it readable by root only:

chmod 600 /etc/netbird/relay/relay.env

Then create /etc/netbird/relay/docker-compose.yml:

services:
  relay:
    image: netbirdio/relay:latest
    container_name: netbird-relay
    restart: unless-stopped
    ports:
      - "443:443/tcp"    # relay over WebSocket
      - "443:443/udp"    # relay over QUIC
      - "3478:3478/udp"  # STUN
    env_file:
      - relay.env
    volumes:
      - relay_data:/data

volumes:
  relay_data:

Repeat on every instance with its own name, for example eu-1.relay.example.com. Open 443/tcp, 443/udp and 3478/udp to the internet on each. No inbound port 80 is needed: the relay proves its domain to Let's Encrypt over 443.

Option 2: geo-DNS

Option 2 keeps everything in option 1 and adds one shared name, relay.example.com, that resolves to a nearby instance. A peer connects to the shared name, lands on an instance, and from then on is known by that instance's own address. Your DNS provider must support location-based or latency-based answers, such as Route 53 geolocation or latency routing.

Keep NB_EXPOSED_ADDRESS unique on every instance, exactly as in option 1. The shared name goes only in server.relays.addresses, never in NB_EXPOSED_ADDRESS.

Two things change:

  • The certificate. A peer's first connection uses the shared name, and its connections to another peer's instance use that instance's name, so every instance must present one certificate valid for both. For example, relay.example.com plus *.relay.example.com. What breaks depends on which name is missing:

    • Without the shared name, no peer can connect at all: x509: certificate is valid for us-1.relay.example.com, not relay.example.com.
    • Without the instance's own name, every peer connects and looks healthy, but peers on different instances cannot relay to each other: x509: certificate is valid for relay.example.com, not us-1.relay.example.com.

    A wildcard needs DNS-based validation from your CA. The relay's built-in Let's Encrypt is not suitable: it proves the domain over port 443, and Let's Encrypt's check of the shared name reaches whichever instance DNS returns to Let's Encrypt. Every other instance fails with acme/autocert: ... no viable challenge type found. Supply the certificate yourself. Delete the three NB_LETSENCRYPT_* lines, set both NB_TLS_CERT_FILE and NB_TLS_KEY_FILE, and mount the files. Restart every instance after each renewal, because the relay reads the files only at startup.

  • Health checks on the DNS record. Peers of a failed instance reconnect through the shared name, so the record must stop returning that instance within seconds. Attach health checks and keep the TTL low, or peers are sent back to the failed instance until the record changes. Once it does, peers re-home by themselves on their next reconnection attempt, which can take a minute or more.

Start and verify each instance

cd /etc/netbird/relay
docker compose up -d
docker compose logs -f relay  # verify startup

Then check each instance by its own name from outside:

curl -v https://us-1.relay.example.com/

A 404 page not found response with SSL certificate verify ok is healthy. With option 2, run the same check against relay.example.com from a client, and confirm the certificate is valid for that name too.

STUN handling

STUN is served by the Relay instances themselves. Each relay container with NB_ENABLE_STUN=true runs an embedded STUN server on UDP/3478 alongside the relay listener. You do not need a separate STUN/TURN service.

Tell peers where to find STUN through server.stuns in the Management configuration, with one entry per Relay instance under that instance's own name.

How Management points at the Relay pool and STUN

Every Management replica's config.yaml lists the Relay pool in server.relays.addresses and STUN in server.stuns. With option 1, list every instance:

server:
  stuns:
    - uri: "stun:us-1.relay.example.com:3478"
      proto: "udp"
    - uri: "stun:eu-1.relay.example.com:3478"
      proto: "udp"

  relays:
    addresses:
      - "rels://us-1.relay.example.com:443"
      - "rels://eu-1.relay.example.com:443"
    secret: "<NB_AUTH_SECRET; identical on every Relay instance>"
    credentialsTTL: "24h"

With option 2, relays.addresses holds only the shared name:

  relays:
    addresses:
      - "rels://relay.example.com:443"

The full Management config.yaml is assembled in Step 7.

Step 6: Deploy the Signal pool

Deploy at least two Signal instances on separate failure domains. Each instance connects to the NATS cluster from Step 3 and uses it to reconcile cross-instance peer signaling. The Signal load balancer can distribute peers across instances freely because NATS reconciles the routing.

The Signal pool requires the enterprise Signal image, which includes the NATS coordination paths needed for active-active replication. The upstream community Signal image (netbirdio/signal) runs in single-node mode only and cannot be used in HA. Use the enterprise Signal image URL provided alongside your license; the compose example below shows the current published path.

Run this on each Signal host:

# /etc/netbird/signal/docker-compose.yml on each Signal host
services:
  signal:
    image: ghcr.io/netbirdio/signal-cloud:latest
    container_name: netbird-signal
    restart: unless-stopped
    ports:
      - "443:443"   # Signal gRPC (registered with the Signal LB)
    environment:
      - NB_LICENSE_KEY=<your enterprise license key>
      - NATS_ENDPOINTS=nats://nats-1.example.com:4222,nats://nats-2.example.com:4222,nats://nats-3.example.com:4222
      - SINGLE_NODE_MODE=false
    command:
      - --log-level
      - info
      - --log-file
      - console
      - --port
      - "443"
Env varRequiredNotes
NB_LICENSE_KEYYesEnterprise license key. Same value on every Signal instance.
NATS_ENDPOINTSYesComma-separated list of NATS cluster endpoints from Step 3. Identical on every Signal instance.
SINGLE_NODE_MODEYesMust be false. When unset or set to true, Signal runs in single-node mode without NATS coordination. This mode is incompatible with HA.

When the NATS URLs carry a password, as in the Step 3 example, set NATS_ENDPOINTS in a file listed under env_file, readable by root only, instead of in docker-compose.yml.

Signal gRPC requires TLS. Either terminate TLS at the Signal load balancer and forward plain HTTP/2 (h2c) to backends on port 443, or pass TLS through to the backends. With TLS pass-through, mount certificates into each Signal container and adjust the --port or listen address as needed.

Start Signal on each host:

docker compose -f /etc/netbird/signal/docker-compose.yml up -d
docker compose logs -f signal

In the logs, look for a NATS connection success message on startup. If you see NATS_ENDPOINTS environment variable not set or a NATS connection error, fix it before continuing. Signal will not work in HA without NATS.

Register the Signal instances in the Signal LB

Add each Signal host to the Signal LB's backend pool on port 443. Once all instances are registered and healthy, verify the LB:

# Quick TCP probe via the LB hostname
nc -zv signal.example.com 443

Cross-instance routing will be verified end-to-end after the Management pool is up and peers connect (Step 8).

How Management points at the Signal LB

Every Management replica's config.yaml references the Signal pool through a single load balancer URL in server.signalUri. The load balancer distributes traffic across Signal instances, so this is always one entry:

server:
  signalUri: "https://signal.example.com:443"

The full Management config.yaml is assembled in Step 7.

Step 7: Configure and deploy the Management pool

Now configure the Management replicas to point at everything you've set up: Postgres, Redis, NATS, the Relay instances, and the Signal LB URL. Every replica runs the Enterprise combined server image, ghcr.io/netbirdio/netbird-server-cloud, which the Enterprise installer also deploys. Distribute the same config.yaml to every replica.

Keep the auth: block from your existing deployment's config.yaml. Without it the server exits at startup with failed to create embedded IDP service: issuer is required.

server.signalUri takes one URL, the Signal load balancer's, which distributes traffic across the Signal instances. server.relays.addresses takes every Relay instance's URL with option 1, or only the geo-DNS name with option 2. See Step 5.

server:
  exposedAddress: "https://netbird.example.com:443"
  dataDir: "/var/lib/netbird/"

  # Embedded identity provider: keep the block from your existing config.yaml
  auth:
    issuer: "https://netbird.example.com/oauth2"
    localAuthDisabled: false
    signKeyRefreshEnabled: false
    sessionCookieEncryptionKey: "<preserved from existing deployment; identical on every replica>"
    dashboardRedirectURIs:
      - "https://netbird.example.com/nb-auth"
      - "https://netbird.example.com/nb-silent-auth"
    cliRedirectURIs:
      - "http://localhost:53000/"

  # External STUN: one entry per Relay instance, see Step 5
  stuns:
    - uri: "stun:us-1.relay.example.com:3478"
      proto: "udp"
    - uri: "stun:eu-1.relay.example.com:3478"
      proto: "udp"

  # Relay pool: every instance (option 1). With option 2, only the geo-DNS name.
  relays:
    addresses:
      - "rels://us-1.relay.example.com:443"
      - "rels://eu-1.relay.example.com:443"
    secret: "<NB_AUTH_SECRET; same value as on every Relay instance>"
    credentialsTTL: "24h"

  # External Signal pool: one URL pointing at the Signal load balancer
  signalUri: "https://signal.example.com:443"

  # HA: enables active-active mode in the Management pool
  ha:
    enabled: true
    natsEndpoints: "nats://nats-1.example.com:4222,nats://nats-2.example.com:4222,nats://nats-3.example.com:4222"
    redisAddr: "redis://redis.example.com:6379"

  # PostgreSQL stores: all three must use PostgreSQL in HA
  store:
    engine: "postgres"
    dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"
    encryptionKey: "<preserved from existing deployment; identical on every replica>"

  activityStore:
    engine: "postgres"
    dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"

  authStore:
    engine: "postgres"
    dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"

  trafficFlow:
    enabled: true
    address: "https://netbird.example.com:443"
    interval: "60s"

Deploy the Management replicas

Bring up replicas one at a time and register each in the Management LB once healthy.

  1. Verify dependencies are reachable from a Management host:
    psql 'host=pg.example.com sslmode=require user=netbird dbname=netbird' -c '\dt'
    redis-cli -u "redis://redis.example.com:6379" ping        # expect: PONG
    nats --server nats://nats-1.example.com:4222 server check connection  # expect: OK
    
  2. Distribute config.yaml to every Management replica host. Verify identical files (sha256sum config.yaml) on each.
  3. Start replica 1. Wait for /oauth2/.well-known/openid-configuration to return 200 and check the logs for Management server created followed by Starting CloudServer.
  4. Register replica 1 in the Management LB. Confirm the dashboard is reachable via https://netbird.example.com/.
  5. Start replica 2, verify health, register in the LB.
  6. Repeat for any additional replicas.

Deploy the flow receiver and enricher

The trafficFlow block above tells peers to send their traffic events to https://netbird.example.com:443. The Management server does not receive them. Two Enterprise services do:

  • The flow receiver (ghcr.io/netbirdio/flow-receiver-cloud) accepts the events from peers over gRPC and publishes them to the traffic-events stream in NATS.
  • The flow enricher (ghcr.io/netbirdio/flow-enricher-cloud) reads the stream and writes the events to PostgreSQL, where the dashboard reads them.

Without them, peers keep connecting normally, but no traffic events ever appear and the client log repeats flow receiver sent no headers.

Run one receiver and one enricher next to each Management replica. If one host fails, peers send their events through the receiver on another host, and the enrichers share the stream without storing an event twice.

services:
  receiver:
    image: ghcr.io/netbirdio/flow-receiver-cloud:latest
    restart: unless-stopped
    ports:
      - "8084:8084"
    environment:
      - NB_LICENSE_KEY=<your enterprise license key>
      - NB_FLOW_LISTEN_PORT=8084
      - NB_FLOW_ADAPTER_TYPE=nats
      - NB_FLOW_NATS_ENDPOINTS=nats://netbird:<client password>@nats-1.example.com:4222,nats://netbird:<client password>@nats-2.example.com:4222,nats://netbird:<client password>@nats-3.example.com:4222
      - NB_FLOW_NATS_STREAM=traffic-events
      - NB_FLOW_NATS_SUBJECT=traffic-events.flow
      - NB_FLOW_AUTH_SECRET=<server.relays.secret from config.yaml>

  enricher:
    image: ghcr.io/netbirdio/flow-enricher-cloud:latest
    restart: unless-stopped
    volumes:
      - netbird_enricher:/var/lib/netbird
    environment:
      - NB_LICENSE_KEY=<your enterprise license key>
      - NB_DATADIR=/var/lib/netbird
      - NB_MANAGEMENT_STORE_ENGINE=postgres
      - NB_MANAGEMENT_POSTGRES_DSN=host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require
      - NETBIRD_STORE_ENGINE_POSTGRES_DSN=host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require
      - NB_TRAFFIC_EVENT_STORE_ENGINE=postgres
      - NB_TRAFFIC_EVENT_POSTGRES_DSN=host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require
      - NB_MANAGEMENT_STORE_KEY=<server.store.encryptionKey from config.yaml>
      - NB_FLOW_ADAPTER_TYPE=nats
      - NB_FLOW_NATS_ENDPOINTS=nats://netbird:<client password>@nats-1.example.com:4222,nats://netbird:<client password>@nats-2.example.com:4222,nats://netbird:<client password>@nats-3.example.com:4222
      - NB_FLOW_NATS_STREAM=traffic-events
      - NB_METRICS_PORT=9091
      - NB_PERSISTENCE_RETENTION_PERIOD=168h

volumes:
  netbird_enricher:

Two values must match what you configured earlier. Each one fails silently when it does not:

SettingMust beIf it is not
NB_FLOW_NATS_SUBJECTA subject under traffic-events., the stream's subjects from Step 3The default, netbird.flow.events, matches no stream. Every publish fails and the receiver logs failed to publish message to nats: nats: no response from stream.
NB_FLOW_AUTH_SECRETThe same value as server.relays.secretManagement signs the peers' flow tokens with the relay secret. The receiver rejects every peer with invalid token validation: invalid signature.

Then route the events to the receivers. On the Management load balancer, send the path prefix /flow.FlowService/ on netbird.example.com to port 8084 on every receiver, over HTTP/2 like the other gRPC paths. Without that route, the events never reach a receiver.

To verify, generate traffic between two peers in a group with traffic events enabled. nats stream info traffic-events shows the message count rising, and GET /api/events/network-traffic returns the events.

Step 8: Verify HA end to end

Run each failure scenario and confirm the expected behavior. Do this in a staging environment first.

ScenarioExpected behavior
Stop one Management replicaThe Management LB marks it unhealthy within seconds; API traffic continues on remaining replicas. Established peer connections (signal, relay) are unaffected because they flow through the Signal LB and directly to the Relay instances, not through the Management LB.
Stop one Signal instanceThe Signal LB marks it unhealthy; peers connected to that instance reconnect via the LB and land on a surviving Signal instance. Cross-peer signaling continues via NATS without interruption to peers that were already connected to a different instance.
Stop one Relay instancePeers homed on that instance move to another by themselves: with option 1 to the next instance that answers, with option 2 to the instance DNS returns once its health check drops the stopped one. Relayed connections through the stopped instance drop and re-establish. Peers homed on other instances are unaffected.
Stop one NATS nodeCluster retains quorum (2 of 3 healthy). Writes to the traffic-events stream continue. Cross-instance Signal routing continues.
Disconnect RedisNew OAuth flows fail until Redis returns; established peer connections continue working. Dynamic log/rate-limit changes pause.
Trigger PostgreSQL failover (managed service)Brief outage during the failover; Management replicas reconnect to the new primary and resume. Existing peer connections survive the gap because they don't touch PostgreSQL on every message.

If any scenario doesn't match the expected behavior, see Troubleshooting below.

Operations

Rolling upgrades

For the Management and Signal pools, drain one instance, upgrade it, return it to the LB, and repeat. Schema migrations on the Management pool run automatically on first startup of any replica; subsequent replicas detect the migrated schema and start without re-running it.

  1. Mark instance #1 as draining in its load balancer. Wait for active gRPC streams or WebSocket connections to drain (or hit your drain timeout).
  2. Stop and re-pull the image on instance #1: docker compose pull && docker compose up -d.
  3. Wait for the instance's health check to pass.
  4. Add instance #1 back to the LB pool.
  5. Repeat for the remaining instances in the pool.

The Relay pool has no load balancer to drain. Upgrade one instance at a time, and peers on it move to another instance by themselves. With option 2, take the instance out of the DNS record first.

Upgrade Relay, Signal, and Management pools in any order. The protocols between pools are stable across patch releases.

Adding or removing instances

To add a Management or Signal instance, provision the host, install the same image with the same configuration as existing instances, start the service, and add it to the load balancer pool once healthy. No additional coordination is required because the new instance reads the same shared state.

A Relay instance is different: it needs its own name, certificate and NB_EXPOSED_ADDRESS. With option 1, add its URL to server.relays.addresses and its STUN entry to server.stuns on every Management replica, then restart the replicas one at a time. With option 2, add it to server.stuns the same way, and to the geo-DNS record once it is healthy.

To remove a Management or Signal instance: mark it as draining in the LB, wait for streams to drain, stop the service. The remaining instances continue serving traffic. To remove a Relay instance, first take it out of server.relays.addresses (option 1) or the geo-DNS record (option 2), and out of server.stuns, then stop it.

Rotating secrets

SecretHow to rotate
POSTGRES_PASSWORD / Postgres user passwordUpdate the password in PostgreSQL, update the DSN in every Management replica's config.yaml, restart Management replicas one at a time.
Redis passwordUpdate Redis, update server.ha.redisAddr on every Management replica, restart Management replicas one at a time.
NB_AUTH_SECRET (relay auth)Update every Relay instance simultaneously, then update server.relays.secret on every Management replica at the same time. Peer relay sessions can be rejected while the values differ. Restart Management and Relay together to minimise this window. On each Relay instance, apply the new value with docker compose up -d: docker compose restart keeps the old one.
server.store.encryptionKeyDo not rotate. This key encrypts data at rest in PostgreSQL. Rotating it makes existing encrypted data unreadable. Plan a fresh deployment if you need to change it.

Troubleshooting

Management replicas fail to decrypt records: cipher: message authentication failed

server.store.encryptionKey differs between Management replicas. Confirm the value is byte-identical on every replica (no trailing newline or quotes added by your secret manager). Restart every replica after fixing.

OAuth flow fails with "invalid PKCE verifier" after failover

The Redis URL changed mid-flow, or Redis was unreachable. Verify the URL resolves to a single stable endpoint from every replica:

redis-cli -u "<server.ha.redisAddr>" ping

Expected: PONG.

Signal logs NATS_ENDPOINTS environment variable not set

SINGLE_NODE_MODE is unset (or set to true) on the Signal instance, or NATS_ENDPOINTS is empty. In HA, every Signal instance must have SINGLE_NODE_MODE=false and a non-empty NATS_ENDPOINTS pointing at the NATS cluster.

Signal messages don't reach peers connected to a different Signal instance

NATS cluster connectivity is broken or NATS_ENDPOINTS is misconfigured on one of the Signal instances. Check the Signal logs on the instance receiving the original peer message. Successful cross-instance routing logs the publish to NATS. If publishing fails, verify the Signal instance can reach every NATS node on port 4222.

Traffic-flow events stop appearing in the dashboard

The NATS cluster lost quorum. Verify two of three nodes are healthy:

nats --server nats://nats-1.example.com:4222 server check jetstream

If a node is down, restart it and wait for it to rejoin the cluster. The traffic-events stream resumes writes as soon as quorum returns.

Relay rejects peer connections: invalid signature

NB_AUTH_SECRET on the Relay instances differs from server.relays.secret on the Management replicas. The relay logs failed to handshake: validate ... invalid signature for every rejected peer. The two values must be byte-identical. Confirm on every Relay container and in every Management replica's config.yaml.

Relayed connections never form between some peers: peer not available

Some pairs of peers stay Connecting while others relay normally, every peer reports the relay as Available, and the client log shows peer not available: ..., context deadline exceeded. The Relay instances share one NB_EXPOSED_ADDRESS, usually because they sit behind a load balancer, so peers on different instances cannot find each other. Give every instance its own address, as in Step 5.

If the client log shows x509: certificate is valid for relay.example.com, not us-1.relay.example.com instead, the addresses are right but an instance's certificate lacks its own name. See option 2 in Step 5.

Management replica refuses to start: Redis

server.ha.enabled: true needs a reachable server.ha.redisAddr. If it is empty, the replica exits with server.ha.redisAddr is required when ha.enabled is true. If Redis cannot be reached, it exits with failed to create shared cache store: dial tcp ...: connect: connection refused. Set the address and confirm connectivity from the replica:

redis-cli -u "<server.ha.redisAddr>" ping

Load balancer health checks all fail simultaneously

All replicas are likely failing to start for the same reason. Common causes are an unreachable PostgreSQL endpoint or an encryption key mismatch. Check the logs on any one replica:

docker compose logs netbird-server | tail -100

Look for failed to connect to postgres, decode encryption key: illegal base64 data, or cipher: message authentication failed.