HA Hub Deployment¶
Status
Multi-instance HA Hub deployment is implemented and documented here. Large-scale performance validation toward the 1000-site target and customer pilots are still in progress, so coordinate with Qvera Support to confirm readiness for your deployment size before going live.
For deployments below ~250 sites, a single well-provisioned Hub host is usually sufficient. HA matters at scale because:
- The Hub becomes a tier-1 operational dependency. When 1000 hospital sites depend on it for management, an outage is a full-day incident.
- A single host's planned maintenance (OS patching, JVM upgrade, hardware replacement) requires every site to lose connectivity for the duration of the restart.
- Cert auto-rotation work is best driven by a leader-elected node; on a single host the rotation worker is also a single point of failure.
Target topology¶
All Hub instances run the same QIE WAR in hub mode, against the same shared database. A load balancer fronts both the dashboard port (typically 8080/443) and the tunnel listener port (typically 8443).
Each remote QIE keeps its single mTLS tunnel to one Hub instance at a time, chosen by whichever instance the LB routes the QIE to on that connection's initial TCP handshake. If that instance dies, the tunnel drops, the QIE reconnects, and the LB places the new tunnel on a surviving instance.
State considerations¶
| State | Where it lives | HA behavior |
|---|---|---|
hub_site, hub_listener_config, hub_registration, hub_client_bundle, hub_site_acl, hub_revoked_cert |
Shared DB | Naturally consistent across instances |
hub_site_status (heartbeat snapshot) |
Shared DB, UPSERTed per heartbeat | Naturally consistent across instances; cross-instance dashboard reads need an instance_id for the "which instance holds this tunnel" question |
hub_audit_entry + hub-audit.log |
DB shared; log file per-host | Investigators must aggregate audit log files from every Hub host. The DB rows are already aggregated |
| Tunnel state (the open WebSocket connection per site) | In-memory per instance | Not migratable. A tunnel is a TCP connection. Failover = drop + reconnect |
| Login throttle counters | Shared DB | DB-backed cluster-wide sliding window keyed by source IP, so the brute-force limit holds regardless of which node a login POST lands on |
| Hub auth / proxy sessions | Shared DB | DB-backed Jetty session store (NullSessionCache + JDBC), so any node serves any request, with no browser stickiness. The browser↔tunnel cross-node gap is bridged by the intra-Hub relay, not by affinity |
Load balancer requirements¶
Dashboard port: no browser stickiness¶
Hub auth sessions live in the shared DB-backed Jetty session store
(NullSessionCache + JDBC), so any node can authenticate any request.
no sticky session affinity is required or wanted. Use plain
round-robin (or least-conns) on the dashboard port.
A browser request usually lands on a node that does not hold the target site's tunnel (~(N−1)/N of requests on an N-node cluster). That browser↔tunnel cross-node gap is bridged by the intra-Hub relay, not by affinity: the front node forwards the request to the node that owns the tunnel and streams the response back. This makes node-to-node relay reachability a hard requirement; see Node connectivity requirements and failure modes below.
Tunnel listener port¶
The tunnel listener port does not need session affinity. Each mTLS tunnel is a separate connection; if one tunnel lands on instance A and another on instance B that is fine. Both instances share the same Hub state via the DB.
Whatever connection-distribution policy your LB defaults to (round-robin, least-conns) is acceptable. Avoid policies that pin specific source IPs to specific backends unless you are using that as a debugging tool.
Readiness probe¶
Set qie.probePort on every node. In hub mode QIE serves a readiness
probe on that port: a plain GET returns 200 when the node is up and
its database is reachable, and 503 otherwise. This is the same option
the engine uses, but the hub implementation reports on main Jetty + DB
reachability rather than the channel manager (which never starts in hub
mode). The port is refused until main Jetty and the Spring context are up,
so a node mid-startup is correctly treated as not-ready.
The probe is active/active and deliberately does not depend on peer
reachability. A partitioned node still returns 200 for its own
browser/proxy traffic, so one node losing sight of its peers does not
cascade the whole cluster out of the LB. Point the LB health check at
qie.probePort and route browser/proxy traffic only to nodes returning
200. There is no PRIMARY_INSTANCE_IP active/standby behavior in hub mode.
Graceful shutdown / drain¶
On shutdown a node drains gracefully so a rolling restart or scale-down does not drop in-flight requests:
- the readiness probe flips to 503, so the load balancer stops sending it new browser/proxy traffic;
- it waits
qie.hubDrainSeconds(default 30) for the LB to notice and for in-flight relayed requests to finish; - it clears its
tunnel_connected_instancerows, so peers take over the sites it owned immediately rather than after the missed-heartbeat abandonment window, and releases the mutexes it held; - its tunnel and relay listeners stop; the sites it held reconnect to a surviving node within their reconnect backoff.
Size the orchestration grace window above the drain. In Kubernetes set
terminationGracePeriodSeconds comfortably greater than qie.hubDrainSeconds
(e.g. 45 to 75s for a 30s drain) so the pod is not force-killed
mid-drain. A preStop hook is not required because the drain runs on
SIGTERM. Abandonment
detection remains the backstop for dirty deaths (crash, kill -9).
Node connectivity requirements and failure modes¶
The cluster is DB-mediated. The shared database is the membership
registry (instance check-in rows, mutexes, the hub event log, and
hub_site_status), plus one direct node-to-node link, the intra-Hub
relay port. A node that starts but cannot participate has almost always
tripped one of these:
Shared database
- Unreachable DB. The node cannot write its instance check-in row or read peer state; its readiness probe should fail it out of the LB.
- Wrong DB. A node pointed at a different database silently forms a
"cluster of one"; it never sees peers because the instance-history table
is the peer list. Every node must use the identical
connection.url.
HA enabled identically on every node
- Hub HA components (instance check-in, the relay, the event log) and the
DB-backed session store activate only when
-Dqie.haEngineis set to a clustered engine, the same flag the engine HA uses. A node missing it runs as an island: in-memory sessions (so operator logins do not carry across nodes) and no relay/event-log. Set-Dqie.mode=huband-Dqie.haEngine=...on every node.
Inter-node relay network
- Relay port blocked between nodes. The relay is the only direct node-to-node link and carries cross-node proxy traffic. It must be reachable node-to-node and must not be published through the external LB (K8s: a ClusterIP service / pod-to-pod with a NetworkPolicy on the relay port; VM: a firewall rule between Hub hosts only). If it is blocked, the node serves the dashboard but ~(N−1)/N of cross-site proxy sessions return 502/504.
-
Wrong advertised relay address. Each node advertises its relay host/port in its instance-history row and peers dial that. If auto-detection picks a loopback / wrong-NIC / pod-internal address, peers cannot reach it. Override with
qie.hubRelayAdvertisedHost/qie.hubRelayAdvertisedPort(a startup WARN fires on a loopback-looking auto-detect).On Kubernetes every pod needs its own advertised host from an identical pod spec, so do not hard-code it. Inject the pod's own IP with the Downward API. The option is read from its environment-variable form (
qie.hubRelayAdvertisedHost→QIE_HUBRELAYADVERTISEDHOST), so the same spec gives each pod a distinct value:env: - name: QIE_HUBRELAYADVERTISEDHOST # each pod advertises its own IP valueFrom: fieldRef: fieldPath: status.podIP - name: QIE_HUBRELAYADVERTISEDPORT # identical for every pod value: "8843" - name: QIE_HUBRELAYSHAREDSECRET # identical for every pod valueFrom: secretKeyRef: name: qie-hub-relay key: secretOnly the host is per-pod; the port and shared secret are the same on every pod. The advertised pod IP is routable pod-to-pod inside the cluster, so keep the relay port off the external LB (see "Relay port blocked between nodes"). A StatefulSet + headless Service is an alternative (inject
metadata.nameand build the per-pod DNS name), butstatus.podIPis simplest for a Deployment. On VMs, set-Dqie.hubRelayAdvertisedHostto each host's own name or IP (never the peer's, because a swap makes each node advertise the other, and cross-node proxying fails with a 502). - Relay shared-secret mismatch. Peers authenticate the relay with a shared secret (qie.hubRelaySharedSecret); the endpoint also accepts the previous value (qie.hubRelaySharedSecretPrevious) so it can be rotated by a rolling config update. A node with a different or missing secret is rejected by peers (HTTP 401). The relay runs HTTP/2 over TLS with a per-node self-signed cert for confidentiality; trust is the shared secret, not the cert, so it is independent of the internal CA rotation. - Relay pool exhausted. Each node's relay listener has its own bounded worker pool (qie.hubRelayMaxWorkers, default 128), separate from main Jetty's, so relay load cannot starve dashboard serving. A long-lived transfer (log tail, large message) holds one worker for its whole duration; when the pool is full the listener returns 503. Size it to peak concurrent cross-node browser requests (operators × ~6), not site count. On the front side the client pools HTTP/2 connections per peer (qie.hubRelayMaxConnectionsPerPeer, default 4), which is small because HTTP/2 multiplexes all concurrent cross-node streams over each connection.
Version & probe
- Version mismatch. A node on a newer build forces older-version nodes to shut down (a new required schema column would error them). Rolling across versions is not supported; a version upgrade is a brief full-Hub restart (sites reconnect after). Config-only rolling updates are fine.
- Mis-wired readiness probe. The Hub probe reports ready on main Jetty + DB reachability (not the engine's channel manager, which never starts in hub mode). A probe wired to wait on the channel manager would mark every hub node unhealthy and the LB would pull the whole cluster; confirm a healthy node returns 200 before relying on it.
Cross-instance fanout¶
Some operations need to take effect on every Hub instance, not just the instance the operator's browser is connected to:
- Cert revocation (every instance must force-close any open tunnel for the revoked site).
- Force-disconnect / takeover signals.
- Listener restart triggered by Restart Listener on the Hub Configuration page.
Fanout rides a hub event log, a
hub_system_event_log table that every node polls on a short interval
(modeled on the engine's SystemEventLog). The operator's UI write
appends an event; each node applies events targeted at it (or broadcast),
and poll position is checkpointed per node. Targeted events name one
instance; broadcast events run everywhere except the originator.
This keeps the DB the only shared coordination state, with no separate pub/sub broker, and makes the relay port the only direct node-to-node link (control fan-out is DB-mediated, not relayed).
Leader-elected operations¶
Some work should run on exactly one instance at a time, not all of them:
- Cert auto-rotation scheduling (the worker that decides "now is the time to send TRUST_REMOVE for cert X").
- Expired-bundle reaping.
- Audit log retention cleanup (when introduced).
The Hub reuses the existing DatabaseSchemaUpdateMutexDAO
pattern (already proven in the codebase for HA engine schema
updates) for leader election. The leader instance acquires the
mutex, runs the work, and releases on shutdown. Other instances do
not contend until the leader releases or expires.
Schema considerations¶
Adding instances to an existing single-instance Hub deployment is a rolling restart:
- Stage the new Hub instance with the same
qie.mode=hubflag and pointed at the existing DB. - Start the new instance. Hibernate schema-update runs once if needed; subsequent instances see the schema already updated.
- Reconfigure the LB to add the new instance.
- Roll the old instance: drain its connections (set its LB weight to 0), wait for sessions to expire, restart.
No re-enrollment is needed. The new instance reads
hub_listener_config from the shared DB and presents the same
server cert. Sites that reconnect to the new instance trust it
because they have always trusted that cert, not the host that
serves it.
Capacity sizing¶
A rough sizing guide for HA deployments (these are starting points to be refined against your own load testing):
| Sites | Hub instances | DB tier |
|---|---|---|
| Up to 250 | 1 (no HA) or 2 (HA) | Modest |
| 250–500 | 2 (HA) | Larger heap, fast IO for heartbeat UPSERTs |
| 500–1000 | 3 (HA) | Same, plus dedicated DB host |
Per-instance JVM heap of 4–8 GB and 4–8 vCPU is a reasonable starting target for 250 sites per instance. The dominant cost at scale is heartbeat UPSERTs and proxied request throughput, not site count by itself.
Sharded vs. HA¶
A regional sharding model (separate Hub per region, each managing its own subset of sites) is technically supported. The architecture makes no assumptions about a single Hub instance, but it is generally not the recommended primary deployment. Reasons:
- Operational burden 4× when you run 4 sharded Hubs (certs, audit logs, upgrade windows, network ingress) per region.
- UX regression: customer admins juggle 4 dashboards instead of one.
- Cross-region visibility lost.
- Does not solve resilience. Each regional Hub still needs HA, so HA + sharding > HA + single Hub.
Where sharding does make sense: hard regulatory or political constraints (e.g. EU data isolation, mandatory in-region operator access). The architecture supports it without code changes, customers can deploy as many independent Hubs as they want; a QIE registers with exactly one Hub.