Troubleshooting¶
This page collects the most common Hub issues and the first diagnostics to run for each. Most problems land in one of three buckets: the tunnel listener is not running, enrollment is failing, or an enrolled tunnel is failing to connect or stay up.
For each symptom, log paths and the relevant audit events are listed; check the file logs first and the audit log second when investigating.
Tunnel listener does not start¶
Symptom: dashboard shows tunnel listener as STOPPED on a configured Hub¶
Open System Administration -> Hub Configuration. The Listener
Status field shows STOPPED, and the Setup Checklist at the top of
the page shows any first-time setup step that is still open.
Check each of these in order:
- Hostname is set. Without a hostname the listener cannot bind a usable server cert. Enter the hostname, click Save, then click Restart Listener.
- Server Cert is set and has a paired private key. If the field is empty, click Generate Server Cert. The SSL Certs / Keys page lists certs; an entry without a key icon cannot be used as a server cert.
- Internal CA Cert is set. If the field is empty, click Initialize Internal CA. Even if the listener could serve TLS without it, the trust manager needs the CA to validate incoming client certs.
-
The configured port is not already in use. On Linux:
If something else holds the port, either change the port on the Hub Configuration page or stop the conflicting process. 5. Logs.
${qie.home}/logs/qie.loghas anERRORfrom theHubTunnelListenerclass indicating which of the above failed. Common patterns:Listener config row missing, listener will not startServer cert ID <uuid> not found in SSL Certs storeInternal CA cert ID <uuid> has no private keyBindException on port <N>
After resolving the underlying cause, click Restart Listener
again. The status field updates to RUNNING on port N once the
listener is bound.
Symptom: listener restarts but immediately stops again¶
Almost always a problem with the server certificate (corrupt, expired, or with a mismatched private key). Symptoms:
- Listener bound briefly then logs an SSL initialization error.
qie.logshows a stack trace fromHubTunnelListenerStartupServicementioningKey store inconsistentor similar.
Generate a fresh server cert with Rotate Server Cert on the Hub Configuration page (it generates, applies, and restarts the listener in one step).
Enrollment failures¶
Symptom: "Bundle has already been consumed"¶
The remote QIE shows this error after the operator clicks Confirm
Registration. The Hub audit log has a
BUNDLE_REUSE_REJECTED event.
| Cause | Resolution |
|---|---|
| Operator already redeemed this bundle on a different QIE | Generate a new bundle for the intended site |
Operator is re-importing the same .qcb file to the same QIE |
Use the Stop Tunnel / Start Tunnel controls on the Register with Hub dialog. You do not need a fresh bundle to restart |
| Attacker intercepted the bundle in transit | Security incident. Revoke the cert issued to the attacker with reason SUSPECTED_COMPROMISE. Investigate the leak. Generate a new bundle and deliver via a more secure channel |
Symptom: "Bundle has expired"¶
Audit event: BUNDLE_EXPIRED_REJECTED. Bundles expire after
Bundle Expiry (days) (default 30). Generate a new bundle.
If your operational workflow routinely exceeds 30 days between bundle generation and import, raise Bundle Expiry (days) in Hub Configuration. Expiry is the only protection against a stale bundle being held indefinitely, so do not raise it indiscriminately. Match it to your actual handover SLAs.
Symptom: "Bundle was cancelled by a Hub operator"¶
Audit event: BUNDLE_REVOKED_REJECTED. Somebody cancelled this bundle
under Pending Bundles before the site
redeemed it, so it can never be redeemed. Cancelling is deliberate, so
find out why before issuing a replacement: the usual reason is that the
bundle went to the wrong recipient.
| Cause | Resolution |
|---|---|
| The bundle was issued to the wrong recipient and recalled | Issue a new bundle for the correct recipient. For a site that is already registered, use Re-enroll rather than Add New Client |
| Two bundles were issued and the wrong one was delivered | Check Pending Bundles for the surviving bundle and deliver that one |
| The bundle was cancelled and nobody knows why | The audit entry records who cancelled it and when. Ask before reissuing |
The revokedBy and revokedAt values on the audit entry name the operator
and the time. A cancellation caught at the moment of commit, because the
site was registering as it was cancelled, records detectedAt: commit.
Symptom: "Subdomain is already registered to another site"¶
Audit event: SITE_REGISTRATION_REJECTED with reason: subdomain-already-registered.
Another site already owns the subdomain label this bundle was issued
for. The bundle is not consumed and can be redeemed once the conflict
clears, so do not issue a replacement first.
| Cause | Resolution |
|---|---|
| The label was issued twice, and the other site registered first | Delete or rename the other site if it was the mistake, or issue a fresh bundle for a different label |
| The site is re-registering after being deleted and recreated | Use Re-enroll on the existing site instead of Add New Client |
| Two registrations raced and this one lost at commit | The detectedAt: commit detail marks this case. The rival registration now owns the label, so a retry is refused the same way: resolve the conflict as in the rows above, then issue a fresh bundle |
Two neighbouring reasons share this event. site-changed-during-reenrollment
means a recovery bundle lost a race against a change to its own site, and
retrying is safe. reenroll-site-not-found means a recovery bundle's target site no longer
exists, so the site must be enrolled fresh. bundle-not-found means the bundle row was gone by the
time the Hub looked; ask for a new bundle.
Symptom: "JWS signature verification failed" (on the QIE side, at import time)¶
The .qcb file is corrupt or was signed by a different Hub.
Common causes:
- File truncation during transfer (size mismatch).
- The bundle was generated by a Hub with a different internal CA. The operator might be holding a bundle from a previous CA.
- Intentional tampering (rare).
Regenerate the bundle from the current Hub. Verify the file size matches what the Hub showed at generation time.
Symptom: "Could not connect to Hub at " during /register¶
The QIE generated the keypair and CSR but cannot reach the Hub to submit them. Check from the QIE host:
# Should return TLS handshake info, not "connection refused" or "timeout"
openssl s_client -connect hub.example.com:8443 -showcerts
Likely causes:
- DNS resolution failure for the Hub hostname.
- Firewall blocking the QIE's egress to the Hub port.
- The Hub's tunnel listener is not running (see above).
The QIE's qie.log reports the underlying socket error.
Symptom: registration succeeds, but tunnel does not connect¶
The remote QIE saved the signed cert successfully, but the Register
with Hub dialog still shows Disconnected for Tunnel Status. The
QIE's qie.log shows mTLS handshake errors.
Likely causes:
| Cause | Diagnostic |
|---|---|
| Hub listener cert and QIE pinned cert mismatch (e.g. Hub rotated cert during enrollment) | Hub audit shows a successful SITE_REGISTERED event but no subsequent TUNNEL_CONNECTED. Solution: regenerate bundle |
| QIE clock skew vs. Hub clock | Both ends log CertificateNotYetValid or CertificateExpired even though the cert just issued. Sync NTP on both hosts |
Wildcard cert covers *.hub.example.com but tunnel listener uses a hostname without the wildcard (e.g. apex hub.example.com) |
The server cert presented on the tunnel port is the Hub's internal-CA-signed cert, not the wildcard. Only the QIE's pinning matters for the tunnel; if the QIE rejects, regenerate bundle |
Sites Dashboard shows red icons¶
Symptom: a site that was green flips to red¶
The tunnel disconnected. This is usually transient.
- Wait 30–60 seconds. Brief red flashes during heartbeat-loss windows are normal (network glitches, brief WAN reconfigurations, main-Jetty restarts on the remote site.
- If the site stays red for several minutes, check the remote QIE's
qie.logfor tunnel reconnect failures. The remote QIE retries on exponential backoff with a 60-second cap. - Audit log: look for the most recent
TUNNEL_DISCONNECTEDfor this site. Thesource_ipand timestamp narrow down what happened.
Common causes:
- Remote QIE was restarted (planned or crash).
- Network path between site and Hub interrupted.
- Site's client cert expired (the Hub-initiated cert cycle could not run, e.g. because the QIE was down through the renewal window).
- Site's client cert was revoked.
Symptom: red flashes cycling rapidly between online and offline¶
If the icon flips green ↔ red every few seconds, especially with
audit log TUNNEL_TAKEOVER_DETECTED events firing repeatedly, you
likely have two QIE installations presenting the same client
cert. The Hub accepts whichever connected most recently, the
old one reconnects, the cycle repeats.
This is a security signal. See
ACL and Audit Log.
Revoke the cert with reason SUSPECTED_COMPROMISE and investigate
how the cert leaked.
Symptom: a site shows connected (green) when it is actually offline¶
The Sites Dashboard derives a site's online state from its tunnel state
and the age of its last heartbeat. A site is shown offline once its
last heartbeat (or its tunnel connect) is older than
qie.hubHeartbeatStaleSeconds (default 45, three missed 15-second
heartbeats), even if the tunnel was never observed closing cleanly.
If a site still shows green when you know it is down:
- Wait out the stale window (default 45 seconds). An ungraceful client
death (power loss,
kill -9, a severed network) leaves no clean tunnel close, so the Hub only learns the site is gone when its heartbeats stop arriving. - A Hub restart reconciles stale connection state on startup, so a site that was connected before the restart shows offline immediately until it actually reconnects. If you still see green right after a restart, confirm the remote QIE did not simply reconnect in the meantime.
- Tune
qie.hubHeartbeatStaleSecondsif the default does not suit your network: lower it to mark sites offline sooner, raise it to tolerate longer heartbeat gaps on a lossy WAN.
Symptom: a site stays disabled even after clicking Enable¶
Check the audit log for a CERT_REVOKED event for this site's cert.
A revoked cert prevents reconnection even when the site is enabled.
Resolution: re-enroll the site with a new bundle. Revoked certs cannot be un-revoked. That is the design.
Proxy issues¶
Symptom: Click a site → new tab opens → "Cannot connect to host"¶
The browser cannot resolve or reach site-<label>.hub.example.com.
-
Wildcard DNS is the most common cause. Verify with:
If this fails, the DNS provider has not propagated the wildcard record, or the wildcard does not exist. 2. Wildcard TLS cert mismatches manifest as browser warnings, not connection failures. Browser address bar shows a cert warning the operator can read. 3. Verify the main Jetty is reachable on the configured port from the operator's workstation.
Symptom: Click a site → tab opens to 502 Bad Gateway¶
The Hub resolved the subdomain but the tunnel went down between the operator's click and the request.
Refresh the dashboard tab; the site row is red. Wait for reconnection or investigate the underlying tunnel failure (above).
Symptom: Click a site → tab opens to 404 Not Found from the Hub (not from the remote QIE)¶
The Hub could not resolve the subdomain label to a
hub_site row. Causes:
- The subdomain label in the URL is misspelled.
- The site was deleted while the operator's tab was already open navigating to it.
- An attacker is probing arbitrary subdomain labels.
404 Not Foundis the response.
Symptom: GWT-RPC errors after logging in to a remote QIE through the proxy¶
If you see ServletConfig has not been initialized or similar
errors in the proxied QIE's UI:
- Confirm whether the Hub is running in path-based mode
(
-Dqie.hubProxyAllowPathBased=true). Subdomain mode does not need any GWT-RPC special handling and should not produce this error. - In path-based mode, the
X-Hub-Proxy-Strip-Pathheader must reach the remote QIE. Some intermediate proxies strip vendor-specific headers; verify the header survives end-to-end.
Subdomain mode is the recommended configuration and should not exhibit these errors.
Client-cert cycling failed¶
Symptom: a site's cert expires unexpectedly¶
Client-cert renewal is Hub-initiated: the Hub's scheduler sends a
CYCLE_CLIENT_CERT_REQUEST over the tunnel starting 30 days before
expiry by default, and the remote QIE responds. (The engine no longer
schedules its own renewal.) If a cert expires anyway:
- Was the QIE connected during the renewal window? Cycling goes over the existing tunnel. A site whose tunnel was down for the full window cannot cycle. See the accepted limitation.
- Check the Hub's audit log for
CLIENT_CERT_CYCLE_STARTED/CLIENT_CERT_CYCLE_FAILEDfor the site, and forRENEW_CSR_REQUEST/RENEW_CSR_RESPONSEactivity around the cert's expiry. - Check the remote QIE's
qie.logforHubCertificateRenewalWorkermessages. It logs the install/cleanup side of each cycle. - As an immediate fix for a still-connected site, open it on the Sites Dashboard → Certificate tab and click Rotate Certificate to force a cycle now.
Manual recovery for an already-expired site is to re-enroll: generate a fresh bundle and import it on the QIE side. The site's identifier UUID is preserved across re-enrollments.
HA cluster issues (M6)¶
These apply only to multi-node Hub deployments (see HA Hub Deployment). The cluster coordinates through the shared database plus the inter-node relay port; most "a node does not join" problems are a break in one of those two links.
Symptom: cross-site proxying returns 502/504 on some requests but not others¶
A request for a site whose tunnel is on a different node must be relayed to that node; on an N-node cluster ~(N−1)/N of requests are cross-node. If roughly that fraction fails while same-node requests succeed, the relay link is broken:
- Confirm the relay port is open node-to-node and is not published through the external LB.
- Confirm each node advertised a reachable relay address (not a loopback /
pod-internal address); override with
qie.hubRelayAdvertisedHost/qie.hubRelayAdvertisedPort. A startup WARN flags a bad auto-detect. - Confirm
qie.hubRelaySharedSecretmatches across nodes (within one rotation step). - The failing request is audited with a "tunnel is on a Hub node this node cannot reach" message naming the owner instance.
Symptom: operator has to log in again when the browser lands on a different node¶
Auth sessions are not being shared. The DB-backed session store activates
only when -Dqie.haEngine=... is set; a node missing it uses in-memory
sessions. Set -Dqie.mode=hub and -Dqie.haEngine=..., identically,
on every node.
Symptom: a node does not see the others / per-node dashboard counts disagree¶
The node is almost certainly on a different database. Every node must
use the identical connection.url. The instance check-in table is the
peer registry, so a node on its own DB forms a silent "cluster of one."
Symptom: the load balancer marks every node unhealthy at once¶
In hub mode the readiness probe reports ready on main-Jetty + DB
reachability (the engine channel manager never starts on a Hub). If all
nodes look unhealthy simultaneously, the probe is mis-wired. Verify a
known-good node returns 200 on qie.probePort directly.
Symptom: a healthy node keeps getting marked stopped / restarting¶
A node that cannot refresh its instance check-in within the abandonment
window (DB slowness, GC pauses, sustained overload) is treated as
abandoned by peers and self-shuts-down. Check qie.log for
instance-checkin failures and the DB's responsiveness under heartbeat
load. Cross-node timing uses DB time, so host clock skew should not
cause this on its own.
Symptom: nodes do not run together after an upgrade¶
A node on a newer build forces older-version nodes to shut down (the schema version guard). Rolling across versions is not supported. Upgrade the whole Hub together (a brief restart; sites reconnect after). Config-only rolling updates are fine.
QIE refuses to start¶
Symptom: startup log shows a dashed block saying the database was set up for the other mode¶
A Hub and an engine each own the database they run against, so QIE checks at every start and stops rather than running against a database that belongs to the other mode. The message is delimited above and below by a line of dashes and lists what was found, for example an installed engine license or a count of configured channels.
QIE stops when it finds:
| Starting as | Refuses when the database holds |
|---|---|
| Hub | an engine license, or any configured channel |
| Engine | a Hub license, any registered site, any issued enrollment bundle, or generated Hub tunnel certificates |
An expired license still counts, because it still shows which mode the database was set up for. A license whose key can no longer be decoded does not.
Do one of the following:
- Point the service at a new, empty database.
- Remove the existing database before starting the service.
- If the mode itself was the mistake, add or remove
-Dqie.mode=huband start again.
The service shuts down rather than restarting, so it will stay down until one of the above is done. See A Hub and an engine cannot share a database.
Symptom: a blank database started in the wrong mode will not switch back¶
It should. Only configuration you created counts, so a database that was started in one mode and never configured is allowed to start in the other. If it is refused, something was configured: check the message for what it names.
Symptom: a non-Hub certificate expired with no warning¶
Certificates other than the Hub's tunnel server leaf and internal CA (the
browser-facing HTTPS certificate, LDAPS or IdP trust certificates) are covered
by the same daily expiry sweep an engine runs, at 90, 30 and 10 days and then
daily under 5. That sweep emails through the alerter, so if nothing arrived,
check that email alerts are configured under System Configuration and look
in qie.log for alerter errors.
The Hub's own two certificates are reported separately, because they are rotated by hand and letting one expire stops every site reconnecting rather than degrading a single integration.
Audit and logs¶
| Want to know about... | Look at |
|---|---|
| Why the listener does not start | ${qie.home}/logs/qie.log for HubTunnelListener errors |
| Why a tunnel does not connect | qie.log on the remote QIE host (for handshake failures) and qie.log on the Hub (for trust manager rejections) |
| Cert-related rejections | Hub audit log (CERT_REVOKED, BUNDLE_REUSE_REJECTED, BUNDLE_EXPIRED_REJECTED, BUNDLE_REVOKED_REJECTED, TUNNEL_TAKEOVER_DETECTED) |
| Operator actions | Hub audit log (SITE_*, ACL_*, BUNDLE_*) |
| Background renewal activity | Hub qie.log for the cycle scheduler and Hub audit CLIENT_CERT_CYCLE_*; remote QIE qie.log for HubCertificateRenewalWorker (install/cleanup side) |
| Proxy request errors | Hub qie.log for ProxyDispatcher and HubProxyFilter |
Getting help¶
When opening a support case for a Hub issue, attach:
- The Hub's
qie.logfrom the affected time window. - The Hub's
hub-audit.logfrom the same window. - The remote QIE's
qie.log(only the affected site's, no need to collect from every site). - A screenshot of System Administration -> Hub Configuration showing the current listener config.
- A description of what the operator saw, and the timestamp of the failure to the second if possible.
These five artifacts allow Qvera Support to reconstruct the failure path without round-tripping for additional logs.