Cluster node that does not start MariaDB after a restart

Prev Next

In a cluster, the database on a node does not come back after a restart. That node stops joining the cluster, and recovery requires Segura® support.

Symptom

  • After restarting the node (or just MariaDB), the service fails within about 30 seconds.

In /var/log/mysql/mysql-error.log:

WSREP: Failed to establish connection: certificate verify failed (SSL routines): certificate has expired
WSREP: Failed to establish connection: sslv3 alert certificate expired (SSL routines)

In systemd:

Failed to start mariadb.service

How to verify

  1. On the failed node, check the database log:
sudo grep "certificate has expired" /var/log/mysql/mysql-error.log | tail -n 5
  1. Access the appliance terminal with a user who has sudo permission. If you do not have this access, open a ticket with Segura® support.

  2. Run

sudo openssl x509 -in /etc/senhasegura/nats/.ssl/ca.crt -noout -enddate
  1. Check the result:
  • notAfter=Oct 1 14:59:33 2026 GMT: the appliance is affected and needs the fix.
  • A date later than 01/10/2026: the fix has already been applied.

Solution

Alert

Do not restart any other node, and do not try to start MariaDB again on the node that went down. Each attempt fails for the same reason and can take down another node in the cluster.

  1. Open a ticket with Segura® support using code SOP 140926. Provide:
    • which nodes are working and which node went down;
    • the time the failed node was restarted;
    • the result of step 4 from the How to verify section on each node;
    • the error lines from the failed node's mysql-error.log;
    • whether the cluster uses a separate arbitrator (garbd) server.
  2. Apply the fix on the working nodes only under support guidance. Support defines the order of steps between the working nodes and the failed node.
  3. After support completes the node recovery, follow "Verify your results" in Segura® maintenance script for on-prem customer.

After the node comes back, its senhasegura-* services may keep restarting for a few minutes with the message The tenant 'senhasegura' is not enabled or does not exist. They recover on their own once the local database finishes starting up.

Messages such as WSREP: Failed to establish connection ... certificate has expired at Note level also appear on the nodes that keep working. They indicate that the cluster is being held up only by connections that were already open.

Do not restore a snapshot of the failed node, and do not reinstall it.

For more information