Skip to content

Troubleshooting

This page collects the most common Hub issues and the first diagnostics to run for each. Most problems land in one of three buckets: the tunnel listener is not running, enrollment is failing, or an enrolled tunnel is failing to connect or stay up.

For each symptom, log paths and the relevant audit events are listed; check the file logs first and the audit log second when investigating.

Tunnel listener does not start

Symptom: dashboard shows tunnel listener as STOPPED on a configured Hub

Open System Administration -> Hub Configuration. The Listener Status field shows STOPPED, and the Setup Checklist at the top of the page shows any first-time setup step that is still open.

Check each of these in order:

  1. Hostname is set. Without a hostname the listener cannot bind a usable server cert. Enter the hostname, click Save, then click Restart Listener.
  2. Server Cert is set and has a paired private key. If the field is empty, click Generate Server Cert. The SSL Certs / Keys page lists certs; an entry without a key icon cannot be used as a server cert.
  3. Internal CA Cert is set. If the field is empty, click Initialize Internal CA. Even if the listener could serve TLS without it, the trust manager needs the CA to validate incoming client certs.
  4. The configured port is not already in use. On Linux:

    ss -lntp | grep -E ':(8443|<your-port>)\b'
    

    If something else holds the port, either change the port on the Hub Configuration page or stop the conflicting process. 5. Logs. ${qie.home}/logs/qie.log has an ERROR from the HubTunnelListener class indicating which of the above failed. Common patterns:

    • Listener config row missing, listener will not start
    • Server cert ID <uuid> not found in SSL Certs store
    • Internal CA cert ID <uuid> has no private key
    • BindException on port <N>

After resolving the underlying cause, click Restart Listener again. The status field updates to RUNNING on port N once the listener is bound.

Symptom: listener restarts but immediately stops again

Almost always a problem with the server certificate (corrupt, expired, or with a mismatched private key). Symptoms:

  • Listener bound briefly then logs an SSL initialization error.
  • qie.log shows a stack trace from HubTunnelListenerStartupService mentioning Key store inconsistent or similar.

Generate a fresh server cert with Rotate Server Cert on the Hub Configuration page (it generates, applies, and restarts the listener in one step).

Enrollment failures

Symptom: "Bundle has already been consumed"

The remote QIE shows this error after the operator clicks Confirm Registration. The Hub audit log has a BUNDLE_REUSE_REJECTED event.

Cause Resolution
Operator already redeemed this bundle on a different QIE Generate a new bundle for the intended site
Operator is re-importing the same .qcb file to the same QIE Use the Stop Tunnel / Start Tunnel controls on the Register with Hub dialog. You do not need a fresh bundle to restart
Attacker intercepted the bundle in transit Security incident. Revoke the cert issued to the attacker with reason SUSPECTED_COMPROMISE. Investigate the leak. Generate a new bundle and deliver via a more secure channel

Symptom: "Bundle has expired"

Audit event: BUNDLE_EXPIRED_REJECTED. Bundles expire after Bundle Expiry (days) (default 30). Generate a new bundle.

If your operational workflow routinely exceeds 30 days between bundle generation and import, raise Bundle Expiry (days) in Hub Configuration. Expiry is the only protection against a stale bundle being held indefinitely, so do not raise it indiscriminately. Match it to your actual handover SLAs.

Symptom: "Bundle was cancelled by a Hub operator"

Audit event: BUNDLE_REVOKED_REJECTED. Somebody cancelled this bundle under Pending Bundles before the site redeemed it, so it can never be redeemed. Cancelling is deliberate, so find out why before issuing a replacement: the usual reason is that the bundle went to the wrong recipient.

Cause Resolution
The bundle was issued to the wrong recipient and recalled Issue a new bundle for the correct recipient. For a site that is already registered, use Re-enroll rather than Add New Client
Two bundles were issued and the wrong one was delivered Check Pending Bundles for the surviving bundle and deliver that one
The bundle was cancelled and nobody knows why The audit entry records who cancelled it and when. Ask before reissuing

The revokedBy and revokedAt values on the audit entry name the operator and the time. A cancellation caught at the moment of commit, because the site was registering as it was cancelled, records detectedAt: commit.

Symptom: "Subdomain is already registered to another site"

Audit event: SITE_REGISTRATION_REJECTED with reason: subdomain-already-registered. Another site already owns the subdomain label this bundle was issued for. The bundle is not consumed and can be redeemed once the conflict clears, so do not issue a replacement first.

Cause Resolution
The label was issued twice, and the other site registered first Delete or rename the other site if it was the mistake, or issue a fresh bundle for a different label
The site is re-registering after being deleted and recreated Use Re-enroll on the existing site instead of Add New Client
Two registrations raced and this one lost at commit The detectedAt: commit detail marks this case. The rival registration now owns the label, so a retry is refused the same way: resolve the conflict as in the rows above, then issue a fresh bundle

Two neighbouring reasons share this event. site-changed-during-reenrollment means a recovery bundle lost a race against a change to its own site, and retrying is safe. reenroll-site-not-found means a recovery bundle's target site no longer exists, so the site must be enrolled fresh. bundle-not-found means the bundle row was gone by the time the Hub looked; ask for a new bundle.

Symptom: "JWS signature verification failed" (on the QIE side, at import time)

The .qcb file is corrupt or was signed by a different Hub. Common causes:

  • File truncation during transfer (size mismatch).
  • The bundle was generated by a Hub with a different internal CA. The operator might be holding a bundle from a previous CA.
  • Intentional tampering (rare).

Regenerate the bundle from the current Hub. Verify the file size matches what the Hub showed at generation time.

Symptom: "Could not connect to Hub at " during /register

The QIE generated the keypair and CSR but cannot reach the Hub to submit them. Check from the QIE host:

# Should return TLS handshake info, not "connection refused" or "timeout"
openssl s_client -connect hub.example.com:8443 -showcerts

Likely causes:

  • DNS resolution failure for the Hub hostname.
  • Firewall blocking the QIE's egress to the Hub port.
  • The Hub's tunnel listener is not running (see above).

The QIE's qie.log reports the underlying socket error.

Symptom: registration succeeds, but tunnel does not connect

The remote QIE saved the signed cert successfully, but the Register with Hub dialog still shows Disconnected for Tunnel Status. The QIE's qie.log shows mTLS handshake errors.

Likely causes:

Cause Diagnostic
Hub listener cert and QIE pinned cert mismatch (e.g. Hub rotated cert during enrollment) Hub audit shows a successful SITE_REGISTERED event but no subsequent TUNNEL_CONNECTED. Solution: regenerate bundle
QIE clock skew vs. Hub clock Both ends log CertificateNotYetValid or CertificateExpired even though the cert just issued. Sync NTP on both hosts
Wildcard cert covers *.hub.example.com but tunnel listener uses a hostname without the wildcard (e.g. apex hub.example.com) The server cert presented on the tunnel port is the Hub's internal-CA-signed cert, not the wildcard. Only the QIE's pinning matters for the tunnel; if the QIE rejects, regenerate bundle

Sites Dashboard shows red icons

Symptom: a site that was green flips to red

The tunnel disconnected. This is usually transient.

  1. Wait 30–60 seconds. Brief red flashes during heartbeat-loss windows are normal (network glitches, brief WAN reconfigurations, main-Jetty restarts on the remote site.
  2. If the site stays red for several minutes, check the remote QIE's qie.log for tunnel reconnect failures. The remote QIE retries on exponential backoff with a 60-second cap.
  3. Audit log: look for the most recent TUNNEL_DISCONNECTED for this site. The source_ip and timestamp narrow down what happened.

Common causes:

  • Remote QIE was restarted (planned or crash).
  • Network path between site and Hub interrupted.
  • Site's client cert expired (the Hub-initiated cert cycle could not run, e.g. because the QIE was down through the renewal window).
  • Site's client cert was revoked.

Symptom: red flashes cycling rapidly between online and offline

If the icon flips green ↔ red every few seconds, especially with audit log TUNNEL_TAKEOVER_DETECTED events firing repeatedly, you likely have two QIE installations presenting the same client cert. The Hub accepts whichever connected most recently, the old one reconnects, the cycle repeats.

This is a security signal. See ACL and Audit Log. Revoke the cert with reason SUSPECTED_COMPROMISE and investigate how the cert leaked.

Symptom: a site shows connected (green) when it is actually offline

The Sites Dashboard derives a site's online state from its tunnel state and the age of its last heartbeat. A site is shown offline once its last heartbeat (or its tunnel connect) is older than qie.hubHeartbeatStaleSeconds (default 45, three missed 15-second heartbeats), even if the tunnel was never observed closing cleanly.

If a site still shows green when you know it is down:

  1. Wait out the stale window (default 45 seconds). An ungraceful client death (power loss, kill -9, a severed network) leaves no clean tunnel close, so the Hub only learns the site is gone when its heartbeats stop arriving.
  2. A Hub restart reconciles stale connection state on startup, so a site that was connected before the restart shows offline immediately until it actually reconnects. If you still see green right after a restart, confirm the remote QIE did not simply reconnect in the meantime.
  3. Tune qie.hubHeartbeatStaleSeconds if the default does not suit your network: lower it to mark sites offline sooner, raise it to tolerate longer heartbeat gaps on a lossy WAN.

Symptom: a site stays disabled even after clicking Enable

Check the audit log for a CERT_REVOKED event for this site's cert. A revoked cert prevents reconnection even when the site is enabled.

Resolution: re-enroll the site with a new bundle. Revoked certs cannot be un-revoked. That is the design.

Proxy issues

Symptom: Click a site → new tab opens → "Cannot connect to host"

The browser cannot resolve or reach site-<label>.hub.example.com.

  1. Wildcard DNS is the most common cause. Verify with:

    nslookup site-foo.hub.example.com
    

    If this fails, the DNS provider has not propagated the wildcard record, or the wildcard does not exist. 2. Wildcard TLS cert mismatches manifest as browser warnings, not connection failures. Browser address bar shows a cert warning the operator can read. 3. Verify the main Jetty is reachable on the configured port from the operator's workstation.

Symptom: Click a site → tab opens to 502 Bad Gateway

The Hub resolved the subdomain but the tunnel went down between the operator's click and the request.

Refresh the dashboard tab; the site row is red. Wait for reconnection or investigate the underlying tunnel failure (above).

Symptom: Click a site → tab opens to 404 Not Found from the Hub (not from the remote QIE)

The Hub could not resolve the subdomain label to a hub_site row. Causes:

  • The subdomain label in the URL is misspelled.
  • The site was deleted while the operator's tab was already open navigating to it.
  • An attacker is probing arbitrary subdomain labels. 404 Not Found is the response.

Symptom: GWT-RPC errors after logging in to a remote QIE through the proxy

If you see ServletConfig has not been initialized or similar errors in the proxied QIE's UI:

  • Confirm whether the Hub is running in path-based mode (-Dqie.hubProxyAllowPathBased=true). Subdomain mode does not need any GWT-RPC special handling and should not produce this error.
  • In path-based mode, the X-Hub-Proxy-Strip-Path header must reach the remote QIE. Some intermediate proxies strip vendor-specific headers; verify the header survives end-to-end.

Subdomain mode is the recommended configuration and should not exhibit these errors.

Client-cert cycling failed

Symptom: a site's cert expires unexpectedly

Client-cert renewal is Hub-initiated: the Hub's scheduler sends a CYCLE_CLIENT_CERT_REQUEST over the tunnel starting 30 days before expiry by default, and the remote QIE responds. (The engine no longer schedules its own renewal.) If a cert expires anyway:

  1. Was the QIE connected during the renewal window? Cycling goes over the existing tunnel. A site whose tunnel was down for the full window cannot cycle. See the accepted limitation.
  2. Check the Hub's audit log for CLIENT_CERT_CYCLE_STARTED / CLIENT_CERT_CYCLE_FAILED for the site, and for RENEW_CSR_REQUEST / RENEW_CSR_RESPONSE activity around the cert's expiry.
  3. Check the remote QIE's qie.log for HubCertificateRenewalWorker messages. It logs the install/cleanup side of each cycle.
  4. As an immediate fix for a still-connected site, open it on the Sites Dashboard → Certificate tab and click Rotate Certificate to force a cycle now.

Manual recovery for an already-expired site is to re-enroll: generate a fresh bundle and import it on the QIE side. The site's identifier UUID is preserved across re-enrollments.

HA cluster issues (M6)

These apply only to multi-node Hub deployments (see HA Hub Deployment). The cluster coordinates through the shared database plus the inter-node relay port; most "a node does not join" problems are a break in one of those two links.

Symptom: cross-site proxying returns 502/504 on some requests but not others

A request for a site whose tunnel is on a different node must be relayed to that node; on an N-node cluster ~(N−1)/N of requests are cross-node. If roughly that fraction fails while same-node requests succeed, the relay link is broken:

  • Confirm the relay port is open node-to-node and is not published through the external LB.
  • Confirm each node advertised a reachable relay address (not a loopback / pod-internal address); override with qie.hubRelayAdvertisedHost / qie.hubRelayAdvertisedPort. A startup WARN flags a bad auto-detect.
  • Confirm qie.hubRelaySharedSecret matches across nodes (within one rotation step).
  • The failing request is audited with a "tunnel is on a Hub node this node cannot reach" message naming the owner instance.

Symptom: operator has to log in again when the browser lands on a different node

Auth sessions are not being shared. The DB-backed session store activates only when -Dqie.haEngine=... is set; a node missing it uses in-memory sessions. Set -Dqie.mode=hub and -Dqie.haEngine=..., identically, on every node.

Symptom: a node does not see the others / per-node dashboard counts disagree

The node is almost certainly on a different database. Every node must use the identical connection.url. The instance check-in table is the peer registry, so a node on its own DB forms a silent "cluster of one."

Symptom: the load balancer marks every node unhealthy at once

In hub mode the readiness probe reports ready on main-Jetty + DB reachability (the engine channel manager never starts on a Hub). If all nodes look unhealthy simultaneously, the probe is mis-wired. Verify a known-good node returns 200 on qie.probePort directly.

Symptom: a healthy node keeps getting marked stopped / restarting

A node that cannot refresh its instance check-in within the abandonment window (DB slowness, GC pauses, sustained overload) is treated as abandoned by peers and self-shuts-down. Check qie.log for instance-checkin failures and the DB's responsiveness under heartbeat load. Cross-node timing uses DB time, so host clock skew should not cause this on its own.

Symptom: nodes do not run together after an upgrade

A node on a newer build forces older-version nodes to shut down (the schema version guard). Rolling across versions is not supported. Upgrade the whole Hub together (a brief restart; sites reconnect after). Config-only rolling updates are fine.

QIE refuses to start

Symptom: startup log shows a dashed block saying the database was set up for the other mode

A Hub and an engine each own the database they run against, so QIE checks at every start and stops rather than running against a database that belongs to the other mode. The message is delimited above and below by a line of dashes and lists what was found, for example an installed engine license or a count of configured channels.

QIE stops when it finds:

Starting as Refuses when the database holds
Hub an engine license, or any configured channel
Engine a Hub license, any registered site, any issued enrollment bundle, or generated Hub tunnel certificates

An expired license still counts, because it still shows which mode the database was set up for. A license whose key can no longer be decoded does not.

Do one of the following:

  1. Point the service at a new, empty database.
  2. Remove the existing database before starting the service.
  3. If the mode itself was the mistake, add or remove -Dqie.mode=hub and start again.

The service shuts down rather than restarting, so it will stay down until one of the above is done. See A Hub and an engine cannot share a database.

Symptom: a blank database started in the wrong mode will not switch back

It should. Only configuration you created counts, so a database that was started in one mode and never configured is allowed to start in the other. If it is refused, something was configured: check the message for what it names.

Symptom: a non-Hub certificate expired with no warning

Certificates other than the Hub's tunnel server leaf and internal CA (the browser-facing HTTPS certificate, LDAPS or IdP trust certificates) are covered by the same daily expiry sweep an engine runs, at 90, 30 and 10 days and then daily under 5. That sweep emails through the alerter, so if nothing arrived, check that email alerts are configured under System Configuration and look in qie.log for alerter errors.

The Hub's own two certificates are reported separately, because they are rotated by hand and letting one expire stops every site reconnecting rather than degrading a single integration.

Audit and logs

Want to know about... Look at
Why the listener does not start ${qie.home}/logs/qie.log for HubTunnelListener errors
Why a tunnel does not connect qie.log on the remote QIE host (for handshake failures) and qie.log on the Hub (for trust manager rejections)
Cert-related rejections Hub audit log (CERT_REVOKED, BUNDLE_REUSE_REJECTED, BUNDLE_EXPIRED_REJECTED, BUNDLE_REVOKED_REJECTED, TUNNEL_TAKEOVER_DETECTED)
Operator actions Hub audit log (SITE_*, ACL_*, BUNDLE_*)
Background renewal activity Hub qie.log for the cycle scheduler and Hub audit CLIENT_CERT_CYCLE_*; remote QIE qie.log for HubCertificateRenewalWorker (install/cleanup side)
Proxy request errors Hub qie.log for ProxyDispatcher and HubProxyFilter

Getting help

When opening a support case for a Hub issue, attach:

  1. The Hub's qie.log from the affected time window.
  2. The Hub's hub-audit.log from the same window.
  3. The remote QIE's qie.log (only the affected site's, no need to collect from every site).
  4. A screenshot of System Administration -> Hub Configuration showing the current listener config.
  5. A description of what the operator saw, and the timestamp of the failure to the second if possible.

These five artifacts allow Qvera Support to reconstruct the failure path without round-tripping for additional logs.