Operator Guide
HA Setup
Commission the supported symmetric two-node pair through the resumable TUI, then prove replication and both failover directions before relying on it.
Prerequisites
- Two fresh Ubuntu 22.04 or 24.04 VPSs with provider-console access and verified SSH root access.
- The signed Server v3.9.18 bootstrap installed on both nodes.
- One final application hostname in a Cloudflare-managed zone.
- A temporary account-scoped token with Workers Scripts Edit, used only while the TUI deploys the witness.
- A long-lived token restricted to Zone Read and DNS Edit for only the application zone. It is stored only as a Worker secret.
- A trusted browser/passkey device and protected off-VPS storage for the recovery identity and exported snapshots.
1. Create Node A and the join code
VPS SSH
mp-opt and choose Fresh two-node HA: create Node A and a join code.- Enter the final application hostname and confirm Node A's public address.
- Enter the temporary Worker deployment token and the zone-scoped DNS token only in the hidden TUI prompts.
- Let the signed commissioning-tools image deploy the witness and bind its exact administrator and DNS secrets.
- Copy the displayed 15-minute Node B join code.
- Leave Node A's TUI open. It polls the witness and continues automatically when Node B completes pairing.
If SSH closes, reconnect and run mp-opt. The protected checkpoint either redisplays the current code or creates its safe replacement.
2. Join Node B
VPS SSH
mp-opt, choose Join an existing HA pair with a one-time code, and paste the code from Node A.The code contains public pairing metadata and a short-lived secret, not application data or a private key. Node B creates its own SSH and age identities; never copy /etc/mp-opt-ha between nodes.
After the witness consumes the code, leave Node B powered on. Node A verifies reciprocal SSH, transfers the exact signed images and complete protected application state, and activates the peer.
3. Accept the first protected copy
Node A captures PostgreSQL, shared configuration, shared Docker secrets, and evidence state into one encrypted bundle. Node B restores into staging, verifies cluster and generation identity, then atomically accepts it.
Server v3.9.18 safely creates a missing empty mount target for the optional node-local Evidence Git token before Compose starts the fresh peer backend. It starts Caddy when the fresh peer has not started it yet, while preserving an already-running instance whose effective configuration has not changed. A configured token is preserved byte-for-byte, and unsafe substitutions are rejected.
Before and after activation, v3.9.18 verifies one runtime permission contract for the HA request queue, result receipts, compliance work, evidence state, generated Caddy policy, service secrets, container mounts and systemd sandboxes. Valid optional-prefixed systemd paths are interpreted correctly; missing, broader or unsafe access still fails closed.
The release also binds witness creation, peer preparation and first-copy acceptance to durable receipts. If SSH or a provider response is lost after the action succeeds, resuming checks those facts before deciding whether anything must be repeated.
First-bundle acceptance and HA service activation are separate checkpoints. Node A and Node B retain matching bundle identity, digest and generation receipts, so reconnecting after an interruption reconciles accepted work instead of transferring the database or recreating healthy services again.
Node B remains at Waiting for first verified copy until PostgreSQL, Backend, and Caddy are all healthy and the receiver has accepted the exact bundle. An established peer keeps its running Caddy instance during ordinary replication when the effective configuration has not changed.
Standby capture runs inside its restricted service identity and reconciles the sender's durable acceptance receipt before the joining peer advances. Scheduled recovery snapshots and setup now share one host-local lease, so a timer catch-up cannot stop the Backend during final validation. A non-holder snapshot invocation exits successfully without creating a competing copy.
4. Finish commissioning and governance
- Complete root commissioning in the browser and publish governance version 1.
- If SMTP is enabled, require successful probes from both origins and matching protected configuration fingerprints.
- Confirm the governance runtime facts declare HA enabled and preview the material notice change before publication.
- Create, deep-verify, and export an independent full recovery snapshot. HA is not a backup.
5. Certify failover before enabling it
- Verify matching release identity, database generation, evidence head, recovery recipient, and recent accepted peer copy.
- Exercise a planned handover from Node A to Node B and back.
- Test narrow synchronous barriers: publisher credentials, public links, and deletion peer confirmation must wait for exact standby proof.
- Exercise automatic failover in both directions using synthetic data and confirm public links and registered passkeys behave as documented.
- Enable automatic failover only after every current gate passes.
Follow the complete verification checklist and routine HA operations.
Safe troubleshooting
- Join code expired or was consumed: resume Node A and use the newly displayed code; do not reuse the old payload.
- Node A says waiting for Node B: leave the TUI open after copying the code. It polls automatically and advances when pairing is verified.
- The deployment completed but public DNS is still pending: leave the TUI open. Version 3.9.18 keeps the verified local deployment checkpoint, compares Cloudflare, Google and Quad9 resolver answers, and retries public routing every 30 seconds. A stale resolver on the VPS does not trigger another deployment.
- The first peer copy is accepted but readiness is still pending: leave the TUI open. The witness may need one short observation interval before it reports the new peer state as healthy. Version 3.9.18 treats that interval as a resumable wait and keeps automatic failover disabled.
- SSH closed during deployment or DNS propagation: reconnect and run
mp-opt. Matching signed images, schema, containers and direct local TLS are reconciled before the TUI resumes the unfinished public-routing step. - Peer rejected the first copy: confirm both management checkouts resolve to the same signed release, then resume. Keep automatic failover disabled.
- Evidence Git is not configured: leave it disabled. The optional empty node-local mount is created automatically.
- Standby protection is unavailable: no new protected mutation is committed when the Backend cannot write the HA queue. Run the installation validation action on the current holder and peer, correct the reported contract failure through the TUI, then retry the original action.
- An existing operation says Attention required: use Retry standby protection. The root-authorised retry reuses the original operation and cannot create a duplicate event, public link, or credential.
- Witness or DNS is unavailable: writes fail closed. Do not force a second writer or edit witness state by hand.
- Node B is waiting for its first verified copy: this is the truthful pre-activation state. Leave it powered on and resume Node A so the holder can transfer, start, and verify the complete peer service.
- Caddy activation or validation failed: keep automatic failover disabled and resume the recorded checkpoint. The TUI reports this boundary separately from Backend health and peer-state finalisation; do not fabricate an acceptance receipt.
- Only Node B's database is running: resume the recorded commissioning checkpoint. The receiver starts a missing Caddy service as part of the verified first-copy activation; do not delete the database or fabricate missing secrets.
Use the immutable v3.9.18 HA guide for the protocol and recovery boundaries.