Problem: A production PostgreSQL 15.17 cluster managed by Patroni 3.3.2, with PgBouncer 1.25.1 in front of it, showed intermittent application failures during night operations. The application logged ODBC connection errors with SQLSTATE 08P01: “server login has been failing, cached error: connect failed (server_login_retry)”. The customer suspected PostgreSQL connection exhaustion. Metrics showed otherwise: max_connections was 4,000 […]
Knowledge Base Data Management and Analytics Database Case Studies Data Management and Analytics Database 18 Sep 2026 PostgreSQL WAL archive failures caused by incomplete copies and full filesystemsProblem: A production PostgreSQL 15 instance experienced a Severity‑1 outage after the filesystem containing WAL files reached 100% usage. The customer observed failed Commvault log-backup jobs and that WAL segments were not being archived: SHOW archive_command returned an external script invocation (/var/lib/pgsql/15/scripts/copy_wal.sh %p %f). The shipped copy script used a “test ! -f … && […]
Knowledge Base Case Studies Data Management and Analytics Database 16 Sep 2026 PostgreSQL replication slot repeatedly went inactive; replica walreceiver stopped and did not reconnectProblem: A Patroni-managed PostgreSQL cluster showed recurring replication interruptions: the replica repeatedly reported a large streaming lag (initially ~565 MB, later observed ~5.5 GB) and the cluster’s physical replication slot on the leader appeared inactive (active = false, wal_status = reserved). Patroni logs on the primary had not recorded recent cluster events and appeared to […]
Knowledge Base Case Studies Data Management and Analytics Database 9 Sep 2026 PostgreSQL SSL handshake failures causing intermittent connection errors in a Patroni clusterProblem: A production PostgreSQL 15.8 deployment running inside a Patroni-managed HA cluster (Patroni 2.1.4) reported transient application failures where processes could not connect and surfaced the message “PostgreSQL connection error, Cannot connect to FX DB”. Database logs showed repeated messages indicating SSL accept-stage failures with the text “could not accept SSL connection: EOF detected”. The […]
Knowledge Base Case Studies Data Management and Analytics Database 9 Sep 2026 PostgreSQL: long-running autovacuum caused by anti-wraparound freeze cycles and index bloatProblem: PostgreSQL 15.17 running in a Patroni-managed cluster experienced autovacuum tasks running for five hours or more on very large, non-partitioned tables. Customer-provided artifacts included table size and index statistics, autovacuum logs, pg_settings, and pg_stat_all_tables. Symptoms reported: multi-hour autovacuum/vacuum runs on large tables, continuous cleanup script execution, and 10-hour REINDEX operations on bloated indexes. The […]
Knowledge Base Case Studies Data Management and Analytics Database 10 Aug 2026 Identifying client IPs and limiting heavy queries for Patroni PostgreSQL behind HAProxyProblem: Production PostgreSQL clusters (PostgreSQL community edition 15.17) are running under Patroni (v3.3.2) with HAProxy (2.6.21) in front. All client connections go through HAProxy using virtual IPs and virtual ports; on the PostgreSQL side the reported client address is the HAProxy VIP rather than the originating physical client IP. The operations team needed a reliable […]
Knowledge Base Case Studies Data Management and Analytics Database 10 Aug 2026 PostgreSQL failover caused by connection surge and memory exhaustion on a Patroni clusterProblem: A production PostgreSQL cluster (community PostgreSQL 15.8, managed by Patroni v2.1.4) experienced an unplanned automatic switchover. Symptoms reported included an unexpected promotion of the standby after the primary stopped responding to health checks. Monitoring and OS traces showed memory usage climbing to ~95% on the primary node, PostgreSQL failing to respond to Patroni probe […]
Case Studies Data Management and Analytics Database 20 May 2026 Resolving a Complex PostgreSQL Patroni Replica FailureProblem The database architecture reached its limit when a critical node stopped responding. A standby node experienced a PostgreSQL Patroni replica failure, repeatedly logging “incorrect resource manager data checksum” errors. The system stopped working because corrupted Write-Ahead Log segments completely broke the replication stream. A dangerous shortcut would involve running continuous base backups to force […]
Database 8 May 2026 Enabling WAL archiving on a DR Patroni standby to allow backups from the replicaProblem: A customer running Patroni-managed PostgreSQL v15.17 (Patroni 3.3.2) asked whether Commvault backups can be taken from the DR site’s Standby Leader. The DR cluster is a replicating standby of Production. The request asked specifically whether WAL file generation can be started on the DR Standby Leader and whether Commvault’s option to delete WALs after […]
Knowledge Base Case Studies Data Management and Analytics Database