PostgreSQL High Availability with Patroni, Etcd, and HAProxy
Building a PostgreSQL High Availability solution with Patroni, Etcd, and HAProxy to eliminate single points of failure and enable automatic failover.

Does your application currently use only a single database server? What happens if your database server suddenly goes down during peak traffic? Your application will face total downtime until the database server is restored. In this scenario, the database becomes a Single Point of Failure.
The most common solution is creating a database replica (Master-Replica), a built-in PostgreSQL feature. However, standard PostgreSQL replication lacks auto-failover, a critical aspect of High Availability (HA) architecture. When the master database dies, we must manually promote a replica to master, update database IPs in our application, and restart the application.
To build a truly High Availability (HA) and self-healing database system, we need three additional components: Patroni, Etcd, and HAProxy.
Patroni, Etcd, and HAProxy
Before diving deeper, let’s get to know these three essential components.
Patroni
Patroni is an open-source tool that acts as a process supervisor to automatically manage the PostgreSQL lifecycle. Patroni runs alongside PostgreSQL on every server node. Its primary functions include:
- Managing initialization, configuration, and database replication.
- Performing periodic database health checks.
- Coordinating with a Distributed Configuration Store (DCS) to elect which node acts as master.
- Providing a REST API to monitor node status (200 for primary, 503 for standby).
Etcd
Etcd is a distributed key-value store acting as a Distributed Configuration Store (DCS). It uses the Raft consensus algorithm, which requires a majority (Quorum) to operate.
- Stores cluster state and metadata (identifying the current Primary and registered Replicas).
- Manages leader locks with a Time-To-Live (TTL) system. The Primary node must regularly renew its lease. If the Primary dies, its key expires, and Patroni on other nodes triggers a leader election.
- Prevents split-brain scenarios (where two nodes believe they are the Primary simultaneously).
We need at least 3 Etcd nodes because the Raft algorithm requires a majority $Q = \lfloor N/2 \rfloor + 1$. In a 2-node cluster, if 1 node dies, the remaining node (50%) does not meet the majority, locking the cluster. With 3 nodes, losing 1 node leaves 2 (66.7%), allowing safe consensus.
HAProxy
HAProxy acts as a high-performance Load Balancer and TCP Proxy sitting in front of the database cluster.
- Provides a single endpoint for the application, so your code doesn’t need to track individual PostgreSQL server IPs.
- Acts as the single source of truth for both reads and writes: all traffic enters via Port
5432, and HAProxy forwards it to the current Primary node. - Checks the Patroni REST API (port
8008). When the Primary dies and a Replica node is promoted, HAProxy automatically detects the new HTTP200 OKresponse and redirects traffic in seconds.
Multi-Server Architecture
To build a truly High Availability cluster and avoid a Single Point of Failure, we will use a 3-database node and 1-gateway topology.
[Mermaid Diagram provided in draft]
Role Distribution
- App Server (Gateway)
- Contains the main application and HAProxy.
- HAProxy acts as the gateway distributing database connections from the application to the correct server.
- Server 1 & 2 (Database Nodes 1 & 2)
- Each runs PostgreSQL, Patroni, and Etcd.
- By default, one is elected Primary (Master) for writes, while others act as Replicas (Standby) receiving automatic data replication.
- Server 3 (Database Node 3 / Quorum Member)
- A full node running PostgreSQL, Patroni, and Etcd.
- Special Role: While capable of becoming a Replica or Primary, its vital function is acting as the tie-breaker (quorum member).
- Etcd needs an odd number of nodes for Quorum ($N/2 + 1$). Server 3 ensures the cluster has 3 Etcd nodes, so if one database server dies, the remaining 2 form a majority to elect a Master automatically.
Hands-on
To simulate a multi-server environment, we use 4 separate Docker compose files connected via a single Docker network.
1docker network create pg-ha-net
[… Node configurations follow here as per the original draft …]
Failover Simulation
The infrastructure is ready, and the application is created; let’s simulate a failover.
Running 3 Node DB Servers and 1 Node HAProxy
We use make up-all to run all nodes. Check the project repository for the source code!

HAProxy Stats Page
We can monitor the nodes via the HAProxy statistics page, specifically focusing on the pg_back section.

Nodes green in the pg_back table are active primary databases, while red nodes are active replicas.
Why? It’s due to option httpchk. The pg_back backend calls Patroni’s /primary endpoint. The node currently serving as primary returns HTTP 200 OK, marked green (UP), and receives all traffic. Replica nodes return 503, and HAProxy marks them red (DOWN).
If you stop the primary container (e.g., docker stop pg-node2), after a few seconds, it turns red, and a promoted replica turns green. Application traffic shifts automatically without restarts or configuration changes.
Running the Application
Ensure Bun is installed, then clone the repository: https://github.com/kykurniawan/pg-ha-poc.
1cd application && bun start

When we kill the primary database (node 2), node 3 is automatically promoted to primary.

During the promotion process, the application might fail to log for a few seconds. This is normal, as promotion takes time; however, it’s better than total connectivity loss.
Error duration can be shortened by reducing ttl and loop_wait in patroni-bootstrap.yml. Beware: smaller values risk false positives (unnecessary failovers caused by minor network hiccups).
Conclusion
We demonstrated how Patroni automatically promotes a replica when the primary goes down, HAProxy reroutes traffic, and the application recovers automatically without manual intervention.
However, those “few seconds” aren’t free. 3 Etcd nodes for a safe quorum triple infrastructure costs, and aggressive ttl/loop_wait values risk unnecessary failovers. High Availability requires balancing these trade-offs according to application needs.
Note: We simulated this on a single server (Docker). In production, these nodes typically reside on physically separate servers or different locations, communicating via secure private networks.
Extra: Why are Read and Write not separated?
Currently, pg_back is the only HAProxy backend, so all traffic is routed to the Primary. Patroni provides a /replica endpoint (opposite of /primary), which we could use for a separate read backend, but we omitted it here to focus on HA failover.
Originally posted at: https://blog.dot.co.id/articles/postgres-high-availability-menggunakan-patroni-etcd-dan-haproxy-1790129940833