Docs / Running the server
High availability
Two Onyx Voice servers and a small witness on a third machine keep the phone system running when one server fails. This chapter explains what fails over and what does not, how to set up the second server and the witness, the Server › Cluster page, switchovers and updates, and what to do when something goes wrong.
How it works
The two servers share one database. Each server's database is either the primary or a standby that copies every change from the primary as it happens. The console, the user portal and phone provisioning run on both servers and always write to the primary.
The phone engine runs on one server only: the active server. That is the one whose database is the primary and which holds the floating address. The floating address is the address phones, carriers, the DNS name and the router's port forwards all point at. It is always on the active server.
When the active server fails, the other one takes over by itself:
- Its database becomes the primary.
- It takes the floating address and tells the network (the switches learn the new location at once).
- It starts the engine with the phones' registrations, so phones do not have to register again and incoming calls reach them straight away.
Calls in progress at that moment drop. New calls work as soon as the engine is up.
Measured on a test cluster with a real desk phone (October 2026):
| What happened | Result |
|---|---|
| Active server switched off | The other server took over in 28.5 s. The phone's registration was restored without the phone registering again. |
| Network cable of the active server cut | It stopped its engine and database after 12.7 s; the other took over at 28.5 s. |
| Planned switchover | 10 s |
| A failed server comes back | It rejoined as the standby by itself in 6 to 14 s. |
What the standby has
- The database: every setting, extension, call record and so on. While the standby is healthy, copying is synchronous: a change counts as saved only once both servers have it. While the standby is down, the primary carries on alone and never waits for it.
- Files: voicemail, call recordings, prompts and hold music, the phone firmware library and the Let's Encrypt certificate. The standby copies them from the active server over HTTPS on port 8444, signed with the cluster's secret: each change within seconds, plus a full comparison when it starts and every 15 minutes. A file still being written (a call being recorded) is copied once it stops changing.
- Engine state: the active server saves the engine's own records (phone registrations, busy lamp subscriptions, queue log-ins and hot desk log-ins) into the database every 5 seconds. The server that takes over loads them into its engine before starting it.
- Blocked addresses: intrusion prevention blocks an address on both servers.
What is lost in a failover: calls in progress, and whatever happened on the failed server in its last few seconds (a voicemail being left, a recording cut off by the failure, the last seconds of engine state).
The witness
When the two servers lose sight of each other, neither can tell whether the other has failed or only the network between them. The witness, a small service on a third machine, settles it.
- The witness hands out a lease: only the server holding it may be the primary.
- It also records which standby has every change (the candidate). Only that standby may take over by itself.
- The active server keeps its lease while it reaches the witness or its standby is copying from it: two out of three. If it has neither for three quarters of the lease time, it fences itself: it stops its engine and database and drops the floating address, so no calls and no changes land on a server that is cut off.
- A standby takes over only after the lease time plus 8 seconds, so a cut-off active server has always stopped first.
The witness needs very little: a small machine or VM installed from the Onyx Voice image. It listens on HTTPS port 8446. It can share a machine with the witness of an Obsidian Mail Server cluster (port 8445); each has its own port, key and state.
Without a witness, or with automatic failover turned off, nothing fails over by itself. The server whose database is the primary keeps the floating address and runs the engine; you can still switch over by hand.
What you need
- Two servers that each meet the requirements for one (see Installing the server), on the same network.
- A fixed address for each server, outside the DHCP pool.
- A floating address in the same network, also outside the DHCP pool. Point the phone system's DNS name, the router's port forwards and anything a carrier sends calls to at this address.
- A third machine for the witness.
- Open ports: between the two servers TCP 5432 (database) and 8444 (file copies), which the setup opens for you when the firewall is on; from both servers to the witness TCP 8446.
Examples below: server 1 at 10.0.0.21, server 2 at 10.0.0.22, the floating address 10.0.0.20/24, the witness at 10.0.0.30.
The servers send calls from the floating address (see The floating address). Allow the whole server network to reach the witness on 8446, not just the two servers' own addresses.
Setting it up
1. Prepare the first server
- Make sure the server has a fixed address: Server › Network. Preparing refuses an address that comes from DHCP.
- Open Server › Cluster. Under Add a second server, press Prepare for a second server and confirm.
This creates the cluster's secret and the account the standby copies with, and sets up the database for copying. The database and the Onyx Voice service restart once; the page reconnects after about a minute. Calls in progress continue, because the engine keeps running.
From a root shell the same is sudo onyx-os cluster.create.
2. Install and join the second server
- On server 1, open Server › Cluster. Under Add the second server, type the address the new server will have (10.0.0.22) and press Create join token.
- Copy the token with Copy token. It is valid for 24 hours and holds the cluster's secrets: treat it like a password.
- Install server 2 from the Onyx Voice image.
- In the first-boot setup, choose Join an existing server (it takes over if that one fails).
- Give it its own host name (for example voip2.example.com). This only names the machine: phones keep using the phone system's name.
- In the network step choose Static address and enter the address you made the token for. The token works only on that address.
- Finish the time zone, administrator password and firewall steps.
- Paste the token when asked. The setup console may not take a paste: if not, sign in over SSH and run
sudo onyx-setupto do it there.
The setup copies the database from server 1 ("Joining: copying the database from the first server..."), becomes the standby and starts copying voicemail, recordings and prompts in the background. The console accounts, tenants and settings come with the database, so the setup does not ask for them. If joining fails, the server's own database is put back and the setup asks for the token again.
You can also join from a shell on server 2 instead of the setup:
sudo onyx-os cluster.join '{"token": "ONYXJOIN1...."}'
If server 2 already has extensions of its own, this refuses unless you add "force": true; its old database is kept aside, not deleted.
Then restart the Onyx Voice service on server 1, so it knows server 2's database address too: Server › System & updates, Services, Restart next to Onyx Voice (control plane, web console), or sudo systemctl restart onyx-voice. Calls in progress continue: this service is not the engine.
3. Set up the witness
Do this after server 2 has joined: the witness token lists the servers that may use the witness.
- Install the witness machine from the Onyx Voice image. In its first-boot setup choose The first server, then set the host name, a fixed address, the time zone and the administrator password; answer No to NAT and cancel the console account and the tenant.
- On server 1, Server › Cluster, under Set up the witness, type the witness machine's address (10.0.0.30) and press Create witness token. Copy it.
- On the witness machine, turn the phone system off (it does not need it) and start the witness:
sudo systemctl disable --now onyx-voice onyx-voice-engine onyx-voice-init
sudo onyx-os witness.setup '{"token": "ONYXJOIN1...."}'
It answers with the witness's address, https://10.0.0.30:8446. When the witness machine's firewall is on, the setup lets the two servers in on 8446. Let the rest of the server network in too, because the active server talks from the floating address:
sudo ufw allow proto tcp from 10.0.0.0/24 to any port 8446
4. Turn on automatic failover
On Server › Cluster, in Automatic failover:
- Tick Fail over automatically.
- Witness:
https://10.0.0.30:8446. - Floating address: the address with its prefix length,
10.0.0.20/24. - Lease (seconds): how long the active server may be silent before the other takes over (plus 8 seconds). The default 12 is right for most networks; it can be 6 to 60.
- Leave A failed server comes back as the standby by itself ticked.
- Press Save.
Both servers pick it up within half a minute. The active server takes the floating address. Check the Failover column: the active server shows Primary and floating address, the other Standby and can take over.
Keeping the address phones use now
If phones, carriers and port forwards already point at server 1's current address, you can make that address the floating address instead of changing them all. Do this outside business hours.
- Server › Network: add a new fixed address for server 1 itself next to the current one, so it has both, for example
10.0.0.20/24, 10.0.0.21/24. - Prepare the cluster from a root shell, naming the new address as the server's own:
sudo onyx-os cluster.create '{"address": "10.0.0.21"}'. (The Prepare for a second server button would use the current address.) - Server › Cluster, Automatic failover: enter the old address as the Floating address (
10.0.0.20/24), leave Fail over automatically unticked for now, and Save. The server now holds the old address as the floating address. - Server › Network: remove the old address, keeping only
10.0.0.21/24. - Continue with step 2 above (join tokens are then made from 10.0.0.21).
Unattended installs
A server or witness can be installed without anyone at its screen. Attach a small disk labelled ONYXSEED (for a VM, a small ISO) holding a file seed.json. On first boot the setup reads it, does everything, and writes what happened to /var/log/onyx-setup.log. For a second server:
{ "hostname": "voip2.example.com",
"network": { "address": "10.0.0.22/24", "gateway": "10.0.0.1", "dns": ["10.0.0.1"] },
"timezone": "America/Chicago",
"admin": { "sshKeys": ["ssh-ed25519 AAAA... [email protected]"] },
"firewall": true,
"join": "ONYXJOIN1...." }
For the witness, use "witness" with the witness token instead of "join"; the setup then turns the phone system off, opens only SSH on its firewall, and starts the witness. admin needs at least sshKeys or a passwordHash (a crypt hash, for example from openssl passwd -6) for the onyxadmin account. With keys and no password hash, onyxadmin may use sudo without a password.
The Cluster page
Server › Cluster (system administrators) refreshes every 10 seconds. Open it at the phone system's name or the floating address: you then reach the active server.
Servers lists both servers:
| Column | What it shows |
|---|---|
| Server | The server's name (this one marks the server you are on) and its address. |
| State | Running, or Not responding. |
| Calls | Active: runs the engine, or Standby. |
| Database | Primary, Standby or Down. |
| Files | On the standby: Up to date with how many files were copied and removed, First copy running, or Copy failing with the reason. On the active server: "the original". |
| Version | The Onyx Voice version it runs. |
| Last heard | When it last reported in. |
A server that is no longer running has a Remove button, to take it off the list.
Automatic failover has the settings above and a Failover column per server: Primary, Standby, Fenced or Database down, floating address on the server that holds it, and on the standby can take over or not in sync yet. The line below says what the server sees: "standby streaming" or "no standby", "holds the lease" or "lease with" the other server, "witness not answering", and whether its engine runs.
Database replication shows each standby connected to the database, its state (streaming when all is well), whether copying is synchronous, and how far behind it is.
Running a cluster
Planned switchover
To take the active server down for maintenance, hand the calls to the other one first:
- On Server › Cluster, press Switch over to (the standby's name) and confirm.
- Calls in progress drop; phones keep their registrations. The switchover takes about 10 seconds.
The button shows only while automatic failover is on and the standby has every change (can take over). From a root shell on the active server: sudo onyx-ha --switchover.
Updates
Update one server at a time: update the standby, switch over to it, then update the other. The Version column shows what each runs. See Updates.
When a server fails
The standby takes over by itself within about half a minute. When the failed server comes back, it rejoins as the standby by itself (with A failed server comes back as the standby by itself ticked) and catches up. If its database cannot be brought back in line, it makes a fresh copy from the new primary.
A server that restarts and cannot reach the witness does not start its database as a primary, so an old primary can never take changes on its own. If the witness is gone for good, let the database start once with sudo onyx-ha --allow-start, then fix or replace the witness.
Taking over by hand
If the active server is gone and the witness is gone too, nothing takes over by itself. On the standby:
sudo onyx-ha --takeover
This works only on a standby that had every change. sudo onyx-ha --takeover --force takes over anyway; changes made while the standby was not in sync may be lost.
The floating address
While a server holds the floating address, it also sends from it: the engine sends SIP and call audio from the address phones and carriers know, and names it in its messages. That is why firewalls, carriers' IP lists and the witness must allow the floating address.
Backups and restores
A backup cannot be restored onto a cluster server: it would replace the cluster's own settings. Restore onto a single server, then add the second server again. See Backup and restore.
What is not possible yet
- A cluster is two servers; a third cannot join.
- There is no button to turn a cluster back into a single server.
Troubleshooting
- Where to look: the Server › Cluster page first. On each server,
sudo onyx-os hashows the failover service's view (role, lease, floating address), andjournalctl -u onyx-ha-arbiterits log. - The standby shows "not in sync yet" for a long time: it has not been copying for 30 seconds in a row. Check Database replication and the network between the servers (port 5432).
- "Copy failing" under Files: the standby cannot fetch files from the active server on port 8444. The reason is shown under the badge.
- "witness not answering": check that the witness machine runs and that port 8446 is open to the whole server network.
- Monitoring blocked after a failover: a monitoring tool that sends anonymous SIP OPTIONS to the floating address many times a second can get its own address blocked by intrusion prevention on the new active server. Monitor
http://<floating address>/healthzinstead (it answers with"engine": "connected"while the engine runs), or put the monitoring machine on Never block on the Security page.