Problem: A production PostgreSQL 15.17 cluster managed by Patroni 3.3.2, with PgBouncer 1.25.1 in front of it, showed intermittent application failures during night operations. The application logged ODBC connection errors with SQLSTATE 08P01: “server login has been failing, cached error: connect failed (server_login_retry)”. The customer suspected PostgreSQL connection exhaustion. Metrics showed otherwise: max_connections was 4,000 […]
Knowledge Base Data Management and Analytics Database Case Studies Data Management and Analytics Database 1 Oct 2026 Recovering Cassandra nodes blocked by corrupted system_traces SSTables and oversized trace partitionsProblem: An on-prem Apache Cassandra 2.2.5 cluster had two of five nodes in its DR datacenter reporting DOWN in nodetool status, while the primary datacenter remained UP. The cluster had recently gone through a storage migration in which data was moved to larger disks by taking nodes offline and copying files. The failing nodes logged […]
Knowledge Base Data Management and Analytics Database Case Studies 25 Sep 2026 Removing manual DR rebuilds after primary-site failover for a Patroni PostgreSQL deploymentProblem: A two-site PostgreSQL 13.4 deployment managed by Patroni experienced replication breakage at the DR site after any failover inside the primary site. The DR standby leader’s primary_conninfo was hardcoded to the former primary node’s address, so when the primary role moved to the other primary-site node the standby kept attempting to connect to the […]
Knowledge Base Data Management and Analytics Database Case Studies 23 Sep 2026 Resolving intermittent ORA-01013 in Debezium Oracle Source connectors caused by LogMiner timeoutsProblem: Debezium Oracle source connectors were intermittently failing with ORA-01013 (“user requested cancel of current operation”) every few hours. Restarting the connector temporarily restored replication. Connector configuration attempts included custom log mining parameters and raising database.query.timeout.ms from the default 600000 ms to 1200000 ms; the Debezium docs warn that setting database.query.timeout.ms to 0 disables the […]
Knowledge Base Case Studies 18 Sep 2026 PostgreSQL WAL archive failures caused by incomplete copies and full filesystemsProblem: A production PostgreSQL 15 instance experienced a Severity‑1 outage after the filesystem containing WAL files reached 100% usage. The customer observed failed Commvault log-backup jobs and that WAL segments were not being archived: SHOW archive_command returned an external script invocation (/var/lib/pgsql/15/scripts/copy_wal.sh %p %f). The shipped copy script used a “test ! -f … && […]
Knowledge Base Case Studies Data Management and Analytics Database 17 Sep 2026 PostgreSQL 15.17 (Patroni) — investigating high RAM usage with 4,000 connectionsProblem: A production Patroni-managed PostgreSQL 15.17 cluster (Patroni 3.3.2) reported persistent high memory consumption on database VMs, frequently exceeding 80% of allocated RAM and once reaching ~95% on a previous 300GB node. The deployment uses max_connections=4000 (cannot be reduced by the application), huge_pages configured as on with approximately 60,000 pages reserved, and vm.overcommit_memory=2 set per […]
Knowledge Base Data Management and Analytics Database Case Studies 16 Sep 2026 Resolving cross-site RabbitMQ cluster-wide stalls caused by a single site failureProblem: A three-site RabbitMQ cluster experienced frequent incidents where a single site would become “stuck” and this state would cause all three sites to stop processing. The customer reported that one node would intermittently become unresponsive for varying reasons and that after that event, publishers and consumers across the cluster were blocked until manual intervention. […]
Case Studies DevOps Application Development 16 Sep 2026 PostgreSQL replication slot repeatedly went inactive; replica walreceiver stopped and did not reconnectProblem: A Patroni-managed PostgreSQL cluster showed recurring replication interruptions: the replica repeatedly reported a large streaming lag (initially ~565 MB, later observed ~5.5 GB) and the cluster’s physical replication slot on the leader appeared inactive (active = false, wal_status = reserved). Patroni logs on the primary had not recorded recent cluster events and appeared to […]
Knowledge Base Case Studies Data Management and Analytics Database 9 Sep 2026 Resolving a three-site RabbitMQ cluster-wide block caused by a stuck nodeProblem: A three-site RabbitMQ cluster experienced frequent incidents where a single site would become “stuck” and that condition then caused the entire three-site cluster to block. Symptoms reported included intermittent node unresponsiveness, client operations timing out or hanging, and cluster-wide stalls that required manual intervention to recover. The environment was a geographically distributed, multi-site RabbitMQ […]
Case Studies DevOps Application Development