Problem: An on-prem Apache Cassandra 2.2.5 cluster had two of five nodes in its DR datacenter reporting DOWN in nodetool status, while the primary datacenter remained UP. The cluster had recently gone through a storage migration in which data was moved to larger disks by taking nodes offline and copying files. The failing nodes logged […]
Knowledge Base Data Management and Analytics Database Case Studies 25 Sep 2026 Removing manual DR rebuilds after primary-site failover for a Patroni PostgreSQL deploymentProblem: A two-site PostgreSQL 13.4 deployment managed by Patroni experienced replication breakage at the DR site after any failover inside the primary site. The DR standby leader’s primary_conninfo was hardcoded to the former primary node’s address, so when the primary role moved to the other primary-site node the standby kept attempting to connect to the […]
Knowledge Base Data Management and Analytics Database Case Studies 17 Sep 2026 PostgreSQL 15.17 (Patroni) — investigating high RAM usage with 4,000 connectionsProblem: A production Patroni-managed PostgreSQL 15.17 cluster (Patroni 3.3.2) reported persistent high memory consumption on database VMs, frequently exceeding 80% of allocated RAM and once reaching ~95% on a previous 300GB node. The deployment uses max_connections=4000 (cannot be reduced by the application), huge_pages configured as on with approximately 60,000 pages reserved, and vm.overcommit_memory=2 set per […]
Knowledge Base Data Management and Analytics Database Case Studies 4 Aug 2026 Troubleshooting high swap usage on an OpenSearch 2.9 nodeProblem: An OpenSearch 2.9 cluster (3 nodes) generated a critical alert: one node reported swap utilization at ~80%, with the OpenSearch process identified as the primary swap consumer. The node’s runtime Java command line showed -Xms10g -Xmx10g and -XX:MaxDirectMemorySize=5368709120 (5GB). The environment reported ~23 GB total RAM. The customer asked how to reduce swap use […]
Knowledge Base Data Management and Analytics Data Analytics 8 Jul 2026 Troubleshooting PostgreSQL performance degradation caused by monitoring queries and configuration settingsProblem: Production Patroni-managed PostgreSQL 15.17 cluster (Patroni 3.3.2) experienced intermittent performance degradation that caused application delays and connection resets. The environment uses a connection pooler (PGBouncer) fronting ~1,600 read-only connections while the primary accepted ~1,600 direct read-write connections. The team reported long query latencies, mass connection timeouts, and large temporary file creation observed in database […]
Knowledge Base Data Management and Analytics Database 5 Jul 2026 Reducing nightly LWLock contention and slowdowns in PostgreSQL 15.17 with Patroni and PgBouncerProblem: A production Patroni-managed PostgreSQL (community 15.17, Patroni 3.3.2) cluster experienced recurrent nightly performance degradation. Symptoms included large numbers of LWLock wait events and high query latency during night operations reported by the application and visible in the database analytical report. Patroni and PostgreSQL logs were provided for analysis. Operational context: PgBouncer was present as […]
Knowledge Base Data Management and Analytics Database 1 Jun 2026 Controlling heavy queries and resource usage on a Patroni PostgreSQL clusterProblem: A production Patroni-managed PostgreSQL 15 cluster experienced periodic heavy queries that threatened availability. An example slow job ran for ~84 seconds and performed a full scan of a 1.8 TB partitioned table (arbor.CDR_DATA) that uses daily partitions starting in early April. Most clients connect through generic application users rather than distinct personal accounts. The […]
Knowledge Base Database Case Studies 22 May 2026 OSSpedia Root cause analysis: PostgreSQL primary crashed from system-wide file-descriptor exhaustionProblem: A production Patroni-managed PostgreSQL cluster (PostgreSQL v15.17, Patroni 3.3.2) experienced a primary process abort with SIGABRT during normal operation. Server logs reported that a server process was terminated by signal 6 (Aborted) and that the failed process was executing a COMMIT when the postmaster began terminating other server processes. Subsequent messages showed PostgreSQL could […]
Data Management and Analytics Database Case Studies 27 Mar 2026 Data Management and Analytics ChromaDB: The Open-Source Memory Layer for Artificial IntelligenceThe rapid evolution of generative artificial intelligence has created a significant need for systems that can store and retrieve information with human-like semantic understanding. ChromaDB has emerged as a pivotal technology in this landscape, acting as a specialized storage layer that allows applications to “remember” and reason over vast amounts of unstructured data. By bridging […]
Database CHR