Running a Highly Available Self-Hosted Deployment
Updated
A single NetBird server works well for many deployments, but production environments may need to remain available when an individual host or service instance fails. A highly available deployment removes those single points of failure by spreading services across separate failure domains and routing traffic to healthy instances.
NetBird Enterprise supports active-active high availability (HA) for Management and Signal. This guide shows how to move an existing self-hosted deployment to that mode: Management, Signal, and Relay run as independent pools: Management and Signal behind load balancers, and Relay with an address of its own for every instance. PostgreSQL, Redis, and NATS provide the shared state that lets a healthy Management or Signal instance continue serving traffic when another instance fails.
Use this guide when you need to tolerate the loss of a Management, Signal, Relay, or NATS node and perform rolling upgrades with minimal disruption. It assumes an existing Enterprise deployment that already uses PostgreSQL. For a distributed deployment without active-active Management and Signal, see Splitting Your Self-Hosted Deployment.
Active-active HA for Management and Signal requires a NetBird Enterprise commercial license. Pricing and license details are on the on-prem pricing page.
Before you begin
This is an infrastructure migration, not an in-place switch. Build and validate the new dependencies before changing the public endpoints used by your existing deployment.
| Confirm | Why it matters |
|---|---|
| You have a recent, restorable PostgreSQL backup. | Database migration or endpoint changes should always have a rollback path. |
You have recorded the current server.store.encryptionKey value. | Every Management replica must use the exact same key to read existing encrypted data. |
| You can deploy instances in separate failure domains. | Two containers on one host do not provide high availability. |
| You have stable DNS names and load balancers for Management and Signal, and a DNS name for every Relay instance. | Peers and the dashboard must keep using stable URLs while backends change. Relay instances are reached by their own names, not through a load balancer. |
| PostgreSQL and Redis can each expose one highly available, read-write endpoint. | Management replicas need a shared database and cache, not per-instance stores. |
| You can test the change in staging first. | Failure testing is part of validating an HA deployment. |
Deployment order
Build the HA deployment in this sequence:
- Make PostgreSQL and Redis highly available.
- Deploy and validate the three-node NATS cluster.
- Create the Management and Signal load-balancer frontends, and the DNS records for every service, including one per Relay instance.
- Deploy the Relay and Signal pools.
- Configure and deploy the Management pool.
- Validate normal operation and then run the failure tests.
Scope and endpoint choices
This guide uses these public service URLs:
| Service | Example URL | Used by |
|---|---|---|
| Management | https://netbird.example.com | Dashboard, Management API, Management gRPC, and OAuth |
| Signal | https://signal.example.com | NetBird peers for signaling |
| Relay | One per instance, e.g. rels://us-1.relay.example.com:443 and rels://eu-1.relay.example.com:443. With the geo-DNS option, also a shared rels://relay.example.com:443. | NetBird peers that need relay connectivity |
The Management URL can remain the URL used by your existing deployment. If you change it, update the dashboard environment to use the new Management API and gRPC endpoints; see Dashboard environment configuration. The dashboard itself is stateless and can be served behind the Management load balancer or separately, as long as browsers can reach the configured Management URL.
Architecture
A highly available deployment has three service pools. They have different connection patterns, so they scale and fail over independently instead of forcing every host to handle every workload.
- Management pool: Enterprise Management replicas serve the dashboard API, Management gRPC, and OAuth2 endpoints. Fully stateless: any replica can serve any request because all durable state lives in PostgreSQL and all short-lived state lives in Redis. Replicas scale horizontally to absorb dashboard traffic, peer sync, and OAuth flows.
- Signal pool: Enterprise Signal replicas serve the Signal gRPC API. Peers connect to one Signal instance through the load balancer. The NATS cluster reconciles cross-instance peer signaling, so any Signal instance can deliver a message to any peer regardless of its connected instance. This makes Signal active-active in the Enterprise build. See How active-active Signal works.
- Relay pool: Relay instances carry traffic for peers that cannot reach each other directly. Unlike Management and Signal, Relay does not go behind a load balancer. Relay instances share no state, so two peers can be relayed to each other only when each can reach the other's instance by that instance's own address. Every instance therefore has its own URL, and a peer that loses its relay moves to another instance by itself. See Step 5.
If you use traffic events, two more services run alongside the Management pool: a flow receiver and a flow enricher. See Deploy the flow receiver and enricher.
The pools depend on the following shared infrastructure:
- NATS cluster: Routes cross-instance Signal messages, distributes dynamic configuration (log level and rate limits), and carries the traffic-flow event stream. Quorum is mandatory. Loss of quorum stalls cross-instance signaling.
- PostgreSQL HA endpoint: Stores durable Management data, including accounts, peers, policies, OAuth state, and integrations. It is operator-managed. NetBird treats it as an opaque connection string and expects failover to happen transparently behind the DSN.
- Redis HA endpoint: Stores ephemeral cache data, including OAuth PKCE verifiers, peer cache data, and dynamic configuration. It is operator-managed. Losing Redis interrupts in-flight OAuth flows but does not break running peer connections.
The Management and Signal pools each use one stable URL backed by load-balanced instances, so adding or removing one of their instances is transparent to peers. The Relay pool is reached by each instance's own address instead.
The two load balancers can be independent load balancers, frontends on one shared load balancer, or managed load-balancer resources. Choose the model that fits your infrastructure. Management and Signal must each be reached through one stable URL with at least two backend instances and a load balancer that can fail over within seconds.
How active-active Signal works
In single-node mode, the Signal service keeps peer connection state in memory and cannot be replicated. In the Enterprise HA build, every Signal instance connects to the NATS cluster and uses it to route signaling messages between peers connected to different instances. When peer A, connected to Signal 1 through the load balancer, needs to reach peer B, connected to Signal 2, Signal 1 publishes the signaling message to NATS. Signal 2's NATS subscription forwards it to peer B. The Signal load balancer can distribute peers across instances freely because NATS reconciles cross-instance routing. A healthy NATS cluster is mandatory.
Prerequisites
- An active NetBird Enterprise commercial license.
- At least 2 enterprise Management instances, on separate failure domains.
- At least 2 enterprise Signal instances, on separate failure domains.
- At least 2 Relay instances, on separate failure domains.
- At least 3 NATS instances for the coordination cluster, on separate failure domains. NATS can colocate with NetBird hosts, but the 3 NATS instances must be on different failure domains.
- A load balancer for the Management and Signal pools. These can be two independent load balancers, two frontends on one shared load balancer, or two managed load-balancer resources, as long as each pool is reachable through a single stable URL. Both require HTTP/2 + gRPC support. The Relay pool does not use a load balancer.
- Public FQDNs for Management and Signal, for example
netbird.example.comandsignal.example.com, each resolving to its load balancer. One more per Relay instance, for exampleus-1.relay.example.comandeu-1.relay.example.com, each resolving to that instance. With the geo-DNS option, also a sharedrelay.example.com. - Permissions to deploy services, mount configuration and secrets, expose network ports, manage DNS records, and register instances with the load balancers in your environment.
Step 1: Make PostgreSQL highly available
NetBird treats PostgreSQL as an opaque connection string. You give it a DSN; NetBird makes no assumptions about replication, failover, or topology behind that DSN. PostgreSQL HA is your responsibility.
What NetBird requires
- A single, stable connection endpoint that survives node failure. NetBird does not implement read replicas, partitioning, or client-side connection pooling against a list of hosts.
- Read-write access on every store DSN. NetBird writes on every store; a read-only replica is not sufficient.
- All three NetBird stores (
server.store,server.activityStore,server.authStore) configured with PostgreSQL DSNs. They can point at the same database/instance or at separate ones, depending on your isolation requirements. - TLS is recommended (
sslmode=requirein the DSN).
Common HA patterns
| Pattern | Notes |
|---|---|
| Managed PostgreSQL (AWS RDS Multi-AZ, GCP Cloud SQL HA, Azure Database for PostgreSQL, Aiven, Neon) | Easiest. The service exposes a single endpoint and handles failover internally. |
| Streaming replication + Patroni or Stolon, fronted by PgBouncer or HAProxy | Self-hosted. Patroni handles primary election; PgBouncer/HAProxy exposes a single virtual endpoint that follows the elected primary. |
PostgreSQL with pg_auto_failover | Lighter-weight alternative to Patroni. |
For the full PostgreSQL high-availability documentation, see PostgreSQL: High Availability, Load Balancing, and Replication.
If your current PostgreSQL is not HA
If your current PostgreSQL deployment is not highly available, move it behind an HA endpoint before adding multiple Management replicas. Use the database platform your organization already trusts, such as a managed PostgreSQL service, a Patroni or Stolon cluster, pg_auto_failover, or another standard HA pattern. NetBird only depends on the endpoint behavior, so the implementation can follow your existing database operations model. Whichever approach you choose, the constraints below apply.
The encryption key configured as server.store.encryptionKey must remain byte-identical before and after
the PostgreSQL migration. Sensitive fields in the database are encrypted with this key; rotating it during a
migration makes existing encrypted data unreadable. Preserve the value from your current deployment and apply
it to every Management replica in the HA setup.
At a minimum, your migration must:
- Take a logically consistent backup of the current PostgreSQL data. For example, use
pg_dump, a snapshot, or a point-in-time copy. - Restore the data into the new HA PostgreSQL setup, preserving schemas and table ownership.
- Update
server.store.dsn,server.activityStore.dsn, andserver.authStore.dsninconfig.yamlto point at the new HA endpoint. - Keep
server.store.encryptionKeyidentical to the value used by the original deployment.
Validate after migration by starting one replica against the new endpoint and confirming it boots cleanly, the dashboard loads, and existing peers reconnect.
How Management points at PostgreSQL
Every Management replica's config.yaml references the PostgreSQL HA endpoint through three store DSNs. All three can use the same database (or separate databases if you prefer isolation between management, activity, and auth stores):
server:
store:
engine: "postgres"
dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"
encryptionKey: "<preserved from existing deployment; identical on every replica>"
activityStore:
engine: "postgres"
dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"
authStore:
engine: "postgres"
dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"
The full config.yaml is assembled in Step 7.
Step 2: Set up a Redis HA endpoint
NetBird uses Redis as a shared cache across the Management pool. Redis does not hold persistent NetBird state. If Redis becomes unavailable, in-flight OAuth flows fail until it returns. Established peer connections continue working because they do not query Redis for every message.
What NetBird requires
NetBird connects to Redis as a single standalone instance using one URL and the standard Redis protocol. The URL must point to a stable endpoint that handles failover internally. NetBird does not negotiate failover itself.
- URL schemes:
redis://orrediss://(TLS). TLS is strongly recommended for any network-traversing traffic. - One URL only. NetBird does not accept a list of hosts and does not perform client-side failover. Whatever sits behind the URL is responsible for failover.
- No cluster-aware or Sentinel-aware behavior. NetBird does not negotiate the Redis Cluster protocol (MOVED redirects, slot-aware routing) or the Redis Sentinel discovery protocol. The endpoint must look like a standard standalone Redis instance from NetBird's side.
What works for HA
| Approach | Works? | Notes |
|---|---|---|
| Managed Redis with a single primary endpoint: AWS ElastiCache for Redis (cluster mode disabled / replication group with replicas), GCP Memorystore for Redis (Standard tier), Azure Cache for Redis (Standard or Premium tier without clustering), Upstash Redis (regional), Redis Enterprise Cloud (active-passive) | Yes | Easiest path. The service exposes one DNS name and handles failover internally. The DNS name follows the new primary automatically. |
Self-hosted Redis Sentinel + TCP proxy (HAProxy with a Sentinel-aware backend selector, Nginx stream module, or a keepalived-managed VIP that follows the Sentinel-elected master) | Yes | Sentinel elects the primary. The TCP proxy exposes one stable address that forwards to the current primary. NetBird connects to the proxy and does not communicate with Sentinel. |
| Redis Cluster (cluster mode enabled, multi-shard) even with a single-VIP front door | No | NetBird keys can hash to different cluster slots. Only keys served by the shard behind the VIP are reachable. Other reads and writes fail with MOVED errors that NetBird cannot follow. |
| Multiple Redis URLs (round-robin DNS, comma-separated list, etc.) | No | NetBird takes a single URL string. Failover must be transparent below the URL, not negotiated client-side. |
Recommended path
- If you run on a major cloud: Use managed Redis in non-cluster / standalone-with-replication mode. Look for "cluster mode disabled", "replication group", or "standard tier with failover" in your provider's UI. This option has the least operational burden because the provider hides failover behind a single DNS name.
- If you self-host: Redis Sentinel + TCP proxy is the right pattern. See the Redis Sentinel documentation for Sentinel setup. For the proxy layer, HAProxy with a Sentinel-aware health check (or keepalived for an IP-level VIP) is the typical choice.
If you accidentally point NetBird at a cluster-mode Redis (e.g. ElastiCache cluster mode enabled with multiple shards), the deployment may appear to work at first because some keys happen to land on the accessible shard. It can then fail intermittently as other keys hash to inaccessible shards. Confirm your managed Redis is in non-cluster mode before pointing NetBird at it.
What gets stored
| Data | TTL / pattern |
|---|---|
| OAuth PKCE verifiers, one-time proxy tokens | 10 minutes |
| Peer cache (public-key → account/peer ID) | 30 minutes + jitter |
| Dynamic log-level and rate-limit configuration | Polled every 5 minutes, pub/sub for instant updates |
If Redis becomes unreachable, in-flight OAuth flows fail. Established peer connections continue working until cache entries expire.
How Management points at Redis
Every Management replica's config.yaml references the Redis HA endpoint through server.ha.redisAddr:
server:
ha:
enabled: true
redisAddr: "redis://redis.example.com:6379"
# natsEndpoints set in Step 3 below
The full config.yaml is assembled in Step 7.
Step 3: Set up a NATS HA cluster
NetBird uses NATS for coordination between Management and Signal instances. Like PostgreSQL and Redis, the NATS deployment is operator-owned: you can use a managed NATS service, a Kubernetes operator, or a self-hosted cluster, as long as it exposes the behavior NetBird requires.
What NetBird requires
- A highly available NATS cluster with at least three nodes across separate failure domains.
- JetStream enabled with persistent storage.
- Quorum available for writes. In a three-node cluster, at least two nodes must be healthy.
- Client endpoints reachable from every Management and Signal instance.
- TLS and authentication appropriate for your network boundary.
- A replicated
traffic-eventsJetStream stream for traffic-flow events.
| NATS port | Purpose | Exposure |
|---|---|---|
4222/tcp | Client connections from NetBird and Signal instances | Reachable from the Management and Signal pools |
6222/tcp | Cluster routes (NATS to NATS) | Reachable between NATS nodes only, if you self-host the cluster |
8222/tcp | Monitoring / health endpoints | Internal observability, if enabled by your NATS deployment |
Common HA patterns
| Pattern | Notes |
|---|---|
| Managed NATS with JetStream | Lowest operational burden. The provider is responsible for node placement, storage, upgrades, and failover. |
| NATS Kubernetes operator | Good fit when NetBird runs in Kubernetes or your platform team already operates stateful services there. Use persistent volumes and spread pods across zones or nodes. |
| Self-hosted NATS cluster | Works well when you already operate VMs or bare-metal services. Place each node in a separate failure domain and back JetStream with persistent storage. |
Do not run all NATS nodes on the same host. The cluster may appear healthy, but that host is still a single point of failure.
Example: a self-hosted cluster with Docker Compose
This example runs one NATS node per host, on three hosts: nats-1.example.com, nats-2.example.com and nats-3.example.com. A NATS node can share a host with a NetBird node or a Signal instance, but no host may run two NATS nodes.
Generate two passwords, one for NetBird and Signal to connect with and one for the NATS nodes to connect to each other. Use hex, so they are safe inside URLs:
openssl rand -hex 24
On each host, create a directory with two files. The first, nats-server.conf, differs per host in server_name and client_advertise:
server_name: nats-1
port: 4222
client_advertise: "nats-1.example.com:4222"
http_port: 8222
jetstream {
store_dir: /data
}
authorization {
user: netbird
password: "<client password>"
}
cluster {
name: netbird
port: 6222
authorization {
user: route
password: "<route password>"
}
routes: [
"nats-route://route:<route password>@nats-1.example.com:6222"
"nats-route://route:<route password>@nats-2.example.com:6222"
"nats-route://route:<route password>@nats-3.example.com:6222"
]
}
client_advertise is the address each node gives clients for reconnecting. Without it, a node in a container advertises its Docker-internal address, which NetBird and Signal cannot reach.
The second file, docker-compose.yml, is the same on every host:
services:
nats:
image: nats:2.14
container_name: nats
restart: unless-stopped
command: ["-c", "/etc/nats/nats-server.conf"]
ports:
- "4222:4222"
- "6222:6222"
- "127.0.0.1:8222:8222"
volumes:
- ./nats-server.conf:/etc/nats/nats-server.conf:ro
- nats-data:/data
volumes:
nats-data:
The configuration holds both passwords, so keep it readable by root only, then start the node:
chmod 600 nats-server.conf
docker compose up -d
Allow port 4222 from the Management and Signal hosts only, and port 6222 between the NATS hosts only. Set these rules on your network firewall or security groups: ports that Docker publishes bypass host firewalls such as UFW and firewalld. The monitoring port 8222 listens on the host itself only.
Once all three nodes run, check each one:
curl -s http://127.0.0.1:8222/healthz
curl -s http://127.0.0.1:8222/jsz | grep -E '"cluster_size"|"leader"'
Each node answers {"status":"ok"}, reports "cluster_size": 3 and names the same leader. Until all three have started, the nodes log Error trying to connect to route for the ones that are not up yet.
NetBird and Signal connect as the netbird user, so the NATS URLs you give them carry the client password: nats://netbird:<client password>@nats-1.example.com:4222. Both write these URLs to their logs at startup, password included. Keep those logs as private as the configuration.
This example does not configure TLS, so the passwords and all NATS traffic cross the network unencrypted. Keep NATS on a private network, or add TLS as described in the NATS documentation.
The nats CLI commands below also need the credentials: add --user netbird --password "<client password>". The server report command needs a system account, which this example does not create; use the curl checks above instead.
Traffic-flow stream
NetBird expects the traffic-events JetStream stream to be available for traffic-flow events. Configure it with:
| Setting | Value |
|---|---|
| Stream name | traffic-events |
| Subjects | traffic-events.> |
| Storage | File-backed persistent storage |
| Replicas | 3 |
| Retention | Limits-based |
| Max age | 168h |
| Discard policy | Old messages first |
If you manage NATS directly, the equivalent nats CLI command is:
nats --server nats://nats-1.example.com:4222 \
stream add traffic-events \
--subjects "traffic-events.>" \
--storage file \
--replicas 3 \
--retention limits \
--max-age 168h \
--discard old \
--defaults
The flow receiver publishes to netbird.flow.events by default, a subject this stream does not match. Step 7 sets it to one under traffic-events.; see Deploy the flow receiver and enricher.
If your NATS platform manages streams declaratively, apply the same settings through that platform instead. If the stream already exists with one replica, update it to three replicas:
nats stream edit traffic-events --replicas 3
Verify the cluster
Use your NATS platform's health checks to confirm the cluster is healthy and JetStream has quorum. If you use the nats CLI, a typical check is:
nats --server nats://nats-1.example.com:4222 server report jetstream
You should see three nodes, one elected as JetStream meta-leader, and the traffic-events stream showing three replicas with all peers healthy.
How Management points at NATS
Every Management replica's config.yaml references the NATS cluster through server.ha.natsEndpoints, a comma-separated list of all cluster endpoints. Each Management and Signal instance connects to one of them and handles failover and reconnection across the rest internally:
server:
ha:
enabled: true
natsEndpoints: "nats://nats-1.example.com:4222,nats://nats-2.example.com:4222,nats://nats-3.example.com:4222"
# redisAddr from Step 2 also goes here
Signal instances reference the same list via the NATS_ENDPOINTS environment variable (see Step 6). The full Management config.yaml is assembled in Step 7.
Why an explicit list and not a single DNS round-robin URL? A single hostname with multiple A/AAAA records (e.g. nats://nats.example.com:4222) works because the underlying client resolves and reconnects, but recovery time depends on DNS TTL and cache. An explicit comma-separated list of node addresses gives immediate visibility of every node and the fastest possible failover when one becomes unreachable. Prefer the explicit list.
Step 4: Configure the load balancers
Use the load balancer your organization already operates. NetBird works with any Layer 7 load balancer that supports HTTP/2 and gRPC, including HAProxy, NGINX, AWS Application Load Balancer, Google Cloud Load Balancing, Azure Application Gateway, and Envoy.
Plan one load-balancer frontend for each of the Management and Signal pools. These can be separate load balancers or two frontends on a shared load balancer. Each frontend needs its own DNS name and backend pool. Configure the frontends and empty backend pools now, then add instances as you deploy the Signal and Management services in Steps 6 and 7. The Relay pool does not go behind a load balancer; see Step 5.
| Pool | Public FQDN | Backend port | Frontend protocol | Health check |
|---|---|---|---|---|
| Signal | e.g. signal.example.com | 443 | HTTPS, HTTP/2, gRPC | TCP/443 (or gRPC health if supported) |
| Management | e.g. netbird.example.com | 443 | HTTPS, HTTP/2, gRPC | GET /oauth2/.well-known/openid-configuration → HTTP 200 |
Common requirements for every pool:
| Requirement | Notes |
|---|---|
| TLS termination | Prefer terminating TLS at the load balancer. Set server.tls to empty on the Management replicas. If your environment requires TLS pass-through, configure TLS on every backend instead. |
| HTTP/2 + gRPC support end-to-end | Required for Management and Signal pools. Make sure your load balancer supports long-lived connections. |
| Connection-level affinity | This is the default for Layer 4 load balancers and Layer 7 load balancers that use HTTP/2 streaming. A long-lived connection stays on one backend for its lifetime. No application-level sticky sessions are required. NATS reconciles cross-instance Signal state, while PostgreSQL and Redis hold Management state. |
| Active health checks | Mark backends out of the pool on failure within seconds, not minutes. |
| Connection draining on rolling upgrade | When you remove a backend, the LB should let in-flight gRPC streams finish before tearing them down. |
| Generous idle timeout | Management and Signal gRPC streams can be long-lived. Set the idle timeout above your peer-sync interval. Ten to 30 minutes is a comfortable range for both. |
For the full set of paths and protocols NetBird exposes to each load balancer, see Configuration Files Reference.
The Management check confirms that a replica is up and serving. It keeps returning 200 while PostgreSQL or Redis is unreachable, so monitor those separately.
If you're using a managed cloud load balancer, configure the equivalent of each row above using your provider's UI or infrastructure as code. If you're using a self-hosted reverse proxy, create one backend pool per service, attach health checks, and enable HTTP/2 or WebSocket support as appropriate for each pool.
Step 5: Deploy the Relay pool
Deploy at least two Relay instances on separate failure domains. The Relay pool uses the upstream Relay image (netbirdio/relay). There is no Enterprise-specific Relay image. Relay does not validate a license. It authenticates incoming connections against the shared NB_AUTH_SECRET.
Relay instances share no state with each other, which is why this pool works differently from Management and Signal. The instance a peer connects to becomes its home relay, and announces its own address to that peer. When two peers are on different instances, each reaches the other's instance directly by that announced address. So every instance needs an NB_EXPOSED_ADDRESS of its own, with a DNS name and a TLS certificate that peers can reach.
Do not put Relay instances behind a load balancer that gives them one shared address. Two peers whose connections land on different instances then cannot be relayed to each other: each waits for the other to come online on its own instance, and the connection never forms. Every peer still reports the relay as Available, so the only symptoms are peers that stay Connecting, with peer not available: ..., context deadline exceeded in the client log.
Choose how peers find an instance:
| Option | server.relays.addresses | Which instance a peer uses | Choose it when |
|---|---|---|---|
| 1. List every instance | Every instance's URL | The first to answer: the client connects to all of them in parallel and keeps the fastest, normally the nearest | By default. There is nothing extra to run. |
| 2. Geo-DNS | One shared name that your DNS provider resolves to a nearby instance | The one DNS returns for the peer's location | You run many instances across regions and do not want every peer racing all of them. |
In both options, a peer that loses its relay moves to another instance by itself, and peers on different instances still reach each other.
Option 1: list every instance
Give each instance its own DNS name and let it obtain its own Let's Encrypt certificate for that name. On us-1, create /etc/netbird/relay/relay.env:
NB_LOG_LEVEL=info
NB_LISTEN_ADDRESS=:443
# This instance's own address. Different on every instance.
NB_EXPOSED_ADDRESS=rels://us-1.relay.example.com:443
NB_AUTH_SECRET=<shared secret, identical on every Relay instance and on the Management pool>
NB_LETSENCRYPT_DOMAINS=us-1.relay.example.com
NB_LETSENCRYPT_EMAIL=admin@example.com
NB_LETSENCRYPT_DATA_DIR=/data/letsencrypt
NB_ENABLE_STUN=true
NB_STUN_PORTS=3478
The file holds the shared secret, so make it readable by root only:
chmod 600 /etc/netbird/relay/relay.env
Then create /etc/netbird/relay/docker-compose.yml:
services:
relay:
image: netbirdio/relay:latest
container_name: netbird-relay
restart: unless-stopped
ports:
- "443:443/tcp" # relay over WebSocket
- "443:443/udp" # relay over QUIC
- "3478:3478/udp" # STUN
env_file:
- relay.env
volumes:
- relay_data:/data
volumes:
relay_data:
Repeat on every instance with its own name, for example eu-1.relay.example.com. Open 443/tcp, 443/udp and 3478/udp to the internet on each. No inbound port 80 is needed: the relay proves its domain to Let's Encrypt over 443.
Option 2: geo-DNS
Option 2 keeps everything in option 1 and adds one shared name, relay.example.com, that resolves to a nearby instance. A peer connects to the shared name, lands on an instance, and from then on is known by that instance's own address. Your DNS provider must support location-based or latency-based answers, such as Route 53 geolocation or latency routing.
Keep NB_EXPOSED_ADDRESS unique on every instance, exactly as in option 1. The shared name goes only in server.relays.addresses, never in NB_EXPOSED_ADDRESS.
Two things change:
-
The certificate. A peer's first connection uses the shared name, and its connections to another peer's instance use that instance's name, so every instance must present one certificate valid for both. For example,
relay.example.complus*.relay.example.com. What breaks depends on which name is missing:- Without the shared name, no peer can connect at all:
x509: certificate is valid for us-1.relay.example.com, not relay.example.com. - Without the instance's own name, every peer connects and looks healthy, but peers on different instances cannot relay to each other:
x509: certificate is valid for relay.example.com, not us-1.relay.example.com.
A wildcard needs DNS-based validation from your CA. The relay's built-in Let's Encrypt is not suitable: it proves the domain over port 443, and Let's Encrypt's check of the shared name reaches whichever instance DNS returns to Let's Encrypt. Every other instance fails with
acme/autocert: ... no viable challenge type found. Supply the certificate yourself. Delete the threeNB_LETSENCRYPT_*lines, set bothNB_TLS_CERT_FILEandNB_TLS_KEY_FILE, and mount the files. Restart every instance after each renewal, because the relay reads the files only at startup. - Without the shared name, no peer can connect at all:
-
Health checks on the DNS record. Peers of a failed instance reconnect through the shared name, so the record must stop returning that instance within seconds. Attach health checks and keep the TTL low, or peers are sent back to the failed instance until the record changes. Once it does, peers re-home by themselves on their next reconnection attempt, which can take a minute or more.
Start and verify each instance
cd /etc/netbird/relay
docker compose up -d
docker compose logs -f relay # verify startup
Then check each instance by its own name from outside:
curl -v https://us-1.relay.example.com/
A 404 page not found response with SSL certificate verify ok is healthy. With option 2, run the same check against relay.example.com from a client, and confirm the certificate is valid for that name too.
STUN handling
STUN is served by the Relay instances themselves. Each relay container with NB_ENABLE_STUN=true runs an embedded STUN server on UDP/3478 alongside the relay listener. You do not need a separate STUN/TURN service.
Tell peers where to find STUN through server.stuns in the Management configuration, with one entry per Relay instance under that instance's own name.
How Management points at the Relay pool and STUN
Every Management replica's config.yaml lists the Relay pool in server.relays.addresses and STUN in server.stuns. With option 1, list every instance:
server:
stuns:
- uri: "stun:us-1.relay.example.com:3478"
proto: "udp"
- uri: "stun:eu-1.relay.example.com:3478"
proto: "udp"
relays:
addresses:
- "rels://us-1.relay.example.com:443"
- "rels://eu-1.relay.example.com:443"
secret: "<NB_AUTH_SECRET; identical on every Relay instance>"
credentialsTTL: "24h"
With option 2, relays.addresses holds only the shared name:
relays:
addresses:
- "rels://relay.example.com:443"
The full Management config.yaml is assembled in Step 7.
Step 6: Deploy the Signal pool
Deploy at least two Signal instances on separate failure domains. Each instance connects to the NATS cluster from Step 3 and uses it to reconcile cross-instance peer signaling. The Signal load balancer can distribute peers across instances freely because NATS reconciles the routing.
The Signal pool requires the enterprise Signal image, which includes the NATS coordination paths needed for active-active replication. The upstream community Signal image (netbirdio/signal) runs in single-node mode only and cannot be used in HA. Use the enterprise Signal image URL provided alongside your license; the compose example below shows the current published path.
Run this on each Signal host:
# /etc/netbird/signal/docker-compose.yml on each Signal host
services:
signal:
image: ghcr.io/netbirdio/signal-cloud:latest
container_name: netbird-signal
restart: unless-stopped
ports:
- "443:443" # Signal gRPC (registered with the Signal LB)
environment:
- NB_LICENSE_KEY=<your enterprise license key>
- NATS_ENDPOINTS=nats://nats-1.example.com:4222,nats://nats-2.example.com:4222,nats://nats-3.example.com:4222
- SINGLE_NODE_MODE=false
command:
- --log-level
- info
- --log-file
- console
- --port
- "443"
| Env var | Required | Notes |
|---|---|---|
NB_LICENSE_KEY | Yes | Enterprise license key. Same value on every Signal instance. |
NATS_ENDPOINTS | Yes | Comma-separated list of NATS cluster endpoints from Step 3. Identical on every Signal instance. |
SINGLE_NODE_MODE | Yes | Must be false. When unset or set to true, Signal runs in single-node mode without NATS coordination. This mode is incompatible with HA. |
When the NATS URLs carry a password, as in the Step 3 example, set NATS_ENDPOINTS in a file listed under env_file, readable by root only, instead of in docker-compose.yml.
Signal gRPC requires TLS. Either terminate TLS at the Signal load balancer and forward plain HTTP/2 (h2c) to backends on port 443, or pass TLS through to the backends. With TLS pass-through, mount certificates into each Signal container and adjust the --port or listen address as needed.
Start Signal on each host:
docker compose -f /etc/netbird/signal/docker-compose.yml up -d
docker compose logs -f signal
In the logs, look for a NATS connection success message on startup. If you see NATS_ENDPOINTS environment variable not set or a NATS connection error, fix it before continuing. Signal will not work in HA without NATS.
Register the Signal instances in the Signal LB
Add each Signal host to the Signal LB's backend pool on port 443. Once all instances are registered and healthy, verify the LB:
# Quick TCP probe via the LB hostname
nc -zv signal.example.com 443
Cross-instance routing will be verified end-to-end after the Management pool is up and peers connect (Step 8).
How Management points at the Signal LB
Every Management replica's config.yaml references the Signal pool through a single load balancer URL in server.signalUri. The load balancer distributes traffic across Signal instances, so this is always one entry:
server:
signalUri: "https://signal.example.com:443"
The full Management config.yaml is assembled in Step 7.
Step 7: Configure and deploy the Management pool
Now configure the Management replicas to point at everything you've set up: Postgres, Redis, NATS, the Relay instances, and the Signal LB URL. Every replica runs the Enterprise combined server image, ghcr.io/netbirdio/netbird-server-cloud, which the Enterprise installer also deploys. Distribute the same config.yaml to every replica.
Keep the auth: block from your existing deployment's config.yaml. Without it the server exits at startup with failed to create embedded IDP service: issuer is required.
server.signalUri takes one URL, the Signal load balancer's, which distributes traffic across the Signal instances. server.relays.addresses takes every Relay instance's URL with option 1, or only the geo-DNS name with option 2. See Step 5.
server:
exposedAddress: "https://netbird.example.com:443"
dataDir: "/var/lib/netbird/"
# Embedded identity provider: keep the block from your existing config.yaml
auth:
issuer: "https://netbird.example.com/oauth2"
localAuthDisabled: false
signKeyRefreshEnabled: false
sessionCookieEncryptionKey: "<preserved from existing deployment; identical on every replica>"
dashboardRedirectURIs:
- "https://netbird.example.com/nb-auth"
- "https://netbird.example.com/nb-silent-auth"
cliRedirectURIs:
- "http://localhost:53000/"
# External STUN: one entry per Relay instance, see Step 5
stuns:
- uri: "stun:us-1.relay.example.com:3478"
proto: "udp"
- uri: "stun:eu-1.relay.example.com:3478"
proto: "udp"
# Relay pool: every instance (option 1). With option 2, only the geo-DNS name.
relays:
addresses:
- "rels://us-1.relay.example.com:443"
- "rels://eu-1.relay.example.com:443"
secret: "<NB_AUTH_SECRET; same value as on every Relay instance>"
credentialsTTL: "24h"
# External Signal pool: one URL pointing at the Signal load balancer
signalUri: "https://signal.example.com:443"
# HA: enables active-active mode in the Management pool
ha:
enabled: true
natsEndpoints: "nats://nats-1.example.com:4222,nats://nats-2.example.com:4222,nats://nats-3.example.com:4222"
redisAddr: "redis://redis.example.com:6379"
# PostgreSQL stores: all three must use PostgreSQL in HA
store:
engine: "postgres"
dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"
encryptionKey: "<preserved from existing deployment; identical on every replica>"
activityStore:
engine: "postgres"
dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"
authStore:
engine: "postgres"
dsn: "host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require"
trafficFlow:
enabled: true
address: "https://netbird.example.com:443"
interval: "60s"
Distribute the exact same config.yaml to every Management replica, including the same encryptionKey.
Sensitive fields in PostgreSQL are encrypted with this key. If replicas use different values, they cannot read
each other's records and the cluster can appear to corrupt data on every failover. Use a secret manager such as
Vault, AWS Secrets Manager, SOPS, or sealed-secrets to keep the value synchronized.
Deploy the Management replicas
Bring up replicas one at a time and register each in the Management LB once healthy.
- Verify dependencies are reachable from a Management host:
psql 'host=pg.example.com sslmode=require user=netbird dbname=netbird' -c '\dt' redis-cli -u "redis://redis.example.com:6379" ping # expect: PONG nats --server nats://nats-1.example.com:4222 server check connection # expect: OK - Distribute
config.yamlto every Management replica host. Verify identical files (sha256sum config.yaml) on each. - Start replica 1. Wait for
/oauth2/.well-known/openid-configurationto return 200 and check the logs forManagement server createdfollowed byStarting CloudServer. - Register replica 1 in the Management LB. Confirm the dashboard is reachable via
https://netbird.example.com/. - Start replica 2, verify health, register in the LB.
- Repeat for any additional replicas.
Deploy the flow receiver and enricher
The trafficFlow block above tells peers to send their traffic events to https://netbird.example.com:443. The Management server does not receive them. Two Enterprise services do:
- The flow receiver (
ghcr.io/netbirdio/flow-receiver-cloud) accepts the events from peers over gRPC and publishes them to thetraffic-eventsstream in NATS. - The flow enricher (
ghcr.io/netbirdio/flow-enricher-cloud) reads the stream and writes the events to PostgreSQL, where the dashboard reads them.
Without them, peers keep connecting normally, but no traffic events ever appear and the client log repeats flow receiver sent no headers.
Run one receiver and one enricher next to each Management replica. If one host fails, peers send their events through the receiver on another host, and the enrichers share the stream without storing an event twice.
services:
receiver:
image: ghcr.io/netbirdio/flow-receiver-cloud:latest
restart: unless-stopped
ports:
- "8084:8084"
environment:
- NB_LICENSE_KEY=<your enterprise license key>
- NB_FLOW_LISTEN_PORT=8084
- NB_FLOW_ADAPTER_TYPE=nats
- NB_FLOW_NATS_ENDPOINTS=nats://netbird:<client password>@nats-1.example.com:4222,nats://netbird:<client password>@nats-2.example.com:4222,nats://netbird:<client password>@nats-3.example.com:4222
- NB_FLOW_NATS_STREAM=traffic-events
- NB_FLOW_NATS_SUBJECT=traffic-events.flow
- NB_FLOW_AUTH_SECRET=<server.relays.secret from config.yaml>
enricher:
image: ghcr.io/netbirdio/flow-enricher-cloud:latest
restart: unless-stopped
volumes:
- netbird_enricher:/var/lib/netbird
environment:
- NB_LICENSE_KEY=<your enterprise license key>
- NB_DATADIR=/var/lib/netbird
- NB_MANAGEMENT_STORE_ENGINE=postgres
- NB_MANAGEMENT_POSTGRES_DSN=host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require
- NETBIRD_STORE_ENGINE_POSTGRES_DSN=host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require
- NB_TRAFFIC_EVENT_STORE_ENGINE=postgres
- NB_TRAFFIC_EVENT_POSTGRES_DSN=host=pg.example.com user=netbird password=*** dbname=netbird port=5432 sslmode=require
- NB_MANAGEMENT_STORE_KEY=<server.store.encryptionKey from config.yaml>
- NB_FLOW_ADAPTER_TYPE=nats
- NB_FLOW_NATS_ENDPOINTS=nats://netbird:<client password>@nats-1.example.com:4222,nats://netbird:<client password>@nats-2.example.com:4222,nats://netbird:<client password>@nats-3.example.com:4222
- NB_FLOW_NATS_STREAM=traffic-events
- NB_METRICS_PORT=9091
- NB_PERSISTENCE_RETENTION_PERIOD=168h
volumes:
netbird_enricher:
Two values must match what you configured earlier. Each one fails silently when it does not:
| Setting | Must be | If it is not |
|---|---|---|
NB_FLOW_NATS_SUBJECT | A subject under traffic-events., the stream's subjects from Step 3 | The default, netbird.flow.events, matches no stream. Every publish fails and the receiver logs failed to publish message to nats: nats: no response from stream. |
NB_FLOW_AUTH_SECRET | The same value as server.relays.secret | Management signs the peers' flow tokens with the relay secret. The receiver rejects every peer with invalid token validation: invalid signature. |
Then route the events to the receivers. On the Management load balancer, send the path prefix /flow.FlowService/ on netbird.example.com to port 8084 on every receiver, over HTTP/2 like the other gRPC paths. Without that route, the events never reach a receiver.
To verify, generate traffic between two peers in a group with traffic events enabled. nats stream info traffic-events shows the message count rising, and GET /api/events/network-traffic returns the events.
Step 8: Verify HA end to end
Run each failure scenario and confirm the expected behavior. Do this in a staging environment first.
| Scenario | Expected behavior |
|---|---|
| Stop one Management replica | The Management LB marks it unhealthy within seconds; API traffic continues on remaining replicas. Established peer connections (signal, relay) are unaffected because they flow through the Signal LB and directly to the Relay instances, not through the Management LB. |
| Stop one Signal instance | The Signal LB marks it unhealthy; peers connected to that instance reconnect via the LB and land on a surviving Signal instance. Cross-peer signaling continues via NATS without interruption to peers that were already connected to a different instance. |
| Stop one Relay instance | Peers homed on that instance move to another by themselves: with option 1 to the next instance that answers, with option 2 to the instance DNS returns once its health check drops the stopped one. Relayed connections through the stopped instance drop and re-establish. Peers homed on other instances are unaffected. |
| Stop one NATS node | Cluster retains quorum (2 of 3 healthy). Writes to the traffic-events stream continue. Cross-instance Signal routing continues. |
| Disconnect Redis | New OAuth flows fail until Redis returns; established peer connections continue working. Dynamic log/rate-limit changes pause. |
| Trigger PostgreSQL failover (managed service) | Brief outage during the failover; Management replicas reconnect to the new primary and resume. Existing peer connections survive the gap because they don't touch PostgreSQL on every message. |
If any scenario doesn't match the expected behavior, see Troubleshooting below.
Operations
Rolling upgrades
For the Management and Signal pools, drain one instance, upgrade it, return it to the LB, and repeat. Schema migrations on the Management pool run automatically on first startup of any replica; subsequent replicas detect the migrated schema and start without re-running it.
- Mark instance #1 as draining in its load balancer. Wait for active gRPC streams or WebSocket connections to drain (or hit your drain timeout).
- Stop and re-pull the image on instance #1:
docker compose pull && docker compose up -d. - Wait for the instance's health check to pass.
- Add instance #1 back to the LB pool.
- Repeat for the remaining instances in the pool.
The Relay pool has no load balancer to drain. Upgrade one instance at a time, and peers on it move to another instance by themselves. With option 2, take the instance out of the DNS record first.
Upgrade Relay, Signal, and Management pools in any order. The protocols between pools are stable across patch releases.
Adding or removing instances
To add a Management or Signal instance, provision the host, install the same image with the same configuration as existing instances, start the service, and add it to the load balancer pool once healthy. No additional coordination is required because the new instance reads the same shared state.
A Relay instance is different: it needs its own name, certificate and NB_EXPOSED_ADDRESS. With option 1, add its URL to server.relays.addresses and its STUN entry to server.stuns on every Management replica, then restart the replicas one at a time. With option 2, add it to server.stuns the same way, and to the geo-DNS record once it is healthy.
To remove a Management or Signal instance: mark it as draining in the LB, wait for streams to drain, stop the service. The remaining instances continue serving traffic. To remove a Relay instance, first take it out of server.relays.addresses (option 1) or the geo-DNS record (option 2), and out of server.stuns, then stop it.
Rotating secrets
| Secret | How to rotate |
|---|---|
POSTGRES_PASSWORD / Postgres user password | Update the password in PostgreSQL, update the DSN in every Management replica's config.yaml, restart Management replicas one at a time. |
| Redis password | Update Redis, update server.ha.redisAddr on every Management replica, restart Management replicas one at a time. |
NB_AUTH_SECRET (relay auth) | Update every Relay instance simultaneously, then update server.relays.secret on every Management replica at the same time. Peer relay sessions can be rejected while the values differ. Restart Management and Relay together to minimise this window. On each Relay instance, apply the new value with docker compose up -d: docker compose restart keeps the old one. |
server.store.encryptionKey | Do not rotate. This key encrypts data at rest in PostgreSQL. Rotating it makes existing encrypted data unreadable. Plan a fresh deployment if you need to change it. |
Troubleshooting
Management replicas fail to decrypt records: cipher: message authentication failed
cipher: message authentication failedserver.store.encryptionKey differs between Management replicas. Confirm the value is byte-identical on every replica (no trailing newline or quotes added by your secret manager). Restart every replica after fixing.
OAuth flow fails with "invalid PKCE verifier" after failover
The Redis URL changed mid-flow, or Redis was unreachable. Verify the URL resolves to a single stable endpoint from every replica:
redis-cli -u "<server.ha.redisAddr>" ping
Expected: PONG.
Signal logs NATS_ENDPOINTS environment variable not set
NATS_ENDPOINTS environment variable not setSINGLE_NODE_MODE is unset (or set to true) on the Signal instance, or NATS_ENDPOINTS is empty. In HA, every Signal instance must have SINGLE_NODE_MODE=false and a non-empty NATS_ENDPOINTS pointing at the NATS cluster.
Signal messages don't reach peers connected to a different Signal instance
NATS cluster connectivity is broken or NATS_ENDPOINTS is misconfigured on one of the Signal instances. Check the Signal logs on the instance receiving the original peer message. Successful cross-instance routing logs the publish to NATS. If publishing fails, verify the Signal instance can reach every NATS node on port 4222.
Traffic-flow events stop appearing in the dashboard
The NATS cluster lost quorum. Verify two of three nodes are healthy:
nats --server nats://nats-1.example.com:4222 server check jetstream
If a node is down, restart it and wait for it to rejoin the cluster. The traffic-events stream resumes writes as soon as quorum returns.
Relay rejects peer connections: invalid signature
invalid signatureNB_AUTH_SECRET on the Relay instances differs from server.relays.secret on the Management replicas. The relay logs failed to handshake: validate ... invalid signature for every rejected peer. The two values must be byte-identical. Confirm on every Relay container and in every Management replica's config.yaml.
Relayed connections never form between some peers: peer not available
peer not availableSome pairs of peers stay Connecting while others relay normally, every peer reports the relay as Available, and the client log shows peer not available: ..., context deadline exceeded. The Relay instances share one NB_EXPOSED_ADDRESS, usually because they sit behind a load balancer, so peers on different instances cannot find each other. Give every instance its own address, as in Step 5.
If the client log shows x509: certificate is valid for relay.example.com, not us-1.relay.example.com instead, the addresses are right but an instance's certificate lacks its own name. See option 2 in Step 5.
Management replica refuses to start: Redis
server.ha.enabled: true needs a reachable server.ha.redisAddr. If it is empty, the replica exits with server.ha.redisAddr is required when ha.enabled is true. If Redis cannot be reached, it exits with failed to create shared cache store: dial tcp ...: connect: connection refused. Set the address and confirm connectivity from the replica:
redis-cli -u "<server.ha.redisAddr>" ping
Load balancer health checks all fail simultaneously
All replicas are likely failing to start for the same reason. Common causes are an unreachable PostgreSQL endpoint or an encryption key mismatch. Check the logs on any one replica:
docker compose logs netbird-server | tail -100
Look for failed to connect to postgres, decode encryption key: illegal base64 data, or cipher: message authentication failed.

