Problem:

An on-prem Apache Cassandra 2.2.5 cluster had two of five nodes in its DR datacenter reporting DOWN in nodetool status, while the primary datacenter remained UP. The cluster had recently gone through a storage migration in which data was moved to larger disks by taking nodes offline and copying files. The failing nodes logged CorruptSSTableException and CorruptBlockException errors during compaction. The first corrupt file found was in system_traces.events, alongside extremely large trace-session partitions (about 884 MB and over 2 GB). Further log analysis showed corrupted SSTables in an application keyspace on both nodes as well. One of these nodes had had a data disk replaced several months earlier, and the other also failed to merge the remote schema because of a Lucene index plugin error. During recovery, corruption appeared on a third DR node, leaving only two of five DR nodes live. The customer had routinely stopped Cassandra with kill -9.

Process:

Step 1: Triage each node’s logs separately

Reviewed nodetool status and the system.log of each affected node and identified which table every corrupt SSTable belonged to. Corruption in Cassandra is local to a node, so each node was assessed on its own evidence: a file named in one node’s log is never deleted on another node, and logs of different dates had to be matched to the right host before any action.

Step 2: Quick check for disposable tracing data

On the first node, the corrupt file was in system_traces.events, which holds only query tracing data. As a fast mitigation to try before a full rebuild, the customer was advised to stop Cassandra, delete that table’s SSTable files (Data, Index, Filter, Summary, Statistics, CompressionInfo, TOC), and disable tracing cluster-wide with nodetool settraceprobability 0, since trace partitions of that size destabilize compaction regardless of corruption. Checking disk health (dmesg, smartctl) before restart was required.

Step 3: Escalate when application data is affected

Once the logs showed corrupt SSTables in an application keyspace on both nodes, deleting files was no longer safe. Repeated corruption across disks, including on a node with a previously replaced disk, pointed to file-system or hardware-level damage. With replication factor 3, the reliable fix was to rebuild the nodes and re-replicate their data from healthy replicas rather than repair individual files.

Step 4: Remove and rebuild the dead nodes

The recommended order was: run nodetool repair on live nodes if possible; remove each dead node with nodetool removenode <host-id> from a healthy node, one at a time; confirm the remaining DR nodes are all UN and repair them; wipe data, commit log and saved caches on the removed hosts; rejoin them one at a time, waiting for each bootstrap to move from UJ to UN; then run nodetool repair and nodetool cleanup on all nodes, one at a time.

Step 5: Adapt the plan when DR capacity dropped

When a third node showed corruption and only two DR nodes remained live, a more cautious in-place rebuild was recommended: wipe the broken node completely, start it with -Dcassandra.replace_address=<node IP> in cassandra-env.sh so it streams its data back, monitor with nodetool netstats and status, then remove the flag before moving to the next node. After all nodes were UN, repair and cleanup on every node and a single nodetool repair -dc <dc_name> would restore consistency. If the replace failed, the fallback was to start the wiped node as a new node and rely on datacenter repair.

Step 6: Explain why corruption “disappeared” and returned

After a graceful restart, one node showed no corruption errors, but they reappeared as soon as repair started. Cassandra 2.2.x does not validate all SSTables at startup, so corrupt blocks surface only when compaction or repair reads them, and a node that hits them stops and shows DN. This confirmed that restarts or deleting single files cannot guarantee a clean node, and that remove-and-rejoin is the only reliable path.

Step 7: Address the likely root cause

The repeated use of kill -9 (SIGKILL) to stop Cassandra was identified as the most likely cause of the corruption, possibly combined with disk issues. The customer was given a graceful shutdown procedure: nodetool drain, then nodetool stopdaemon; if the process is still running, kill <PID>; kill -9 only as a last resort. Checking syslog for disk errors and running a full hardware scan, with immediate replacement of any failing disk, was also advised.

Solution:

The recommended remediation for Apache Cassandra was: (1) triage each node’s logs and treat disposable system_traces corruption separately from application data corruption, (2) disable tracing cluster-wide, (3) for nodes with corrupted application SSTables, remove them from the ring or rebuild them in place with replace_address after wiping all data, one node at a time, (4) run nodetool repair and cleanup on all nodes once every node is UN, and (5) stop using kill -9 and switch to a drain-and-stopdaemon shutdown. This works because replication factor 3 keeps healthy copies of the data on other nodes, rebuilding guarantees that every corrupt file is physically gone, and repair brings all replicas back to a consistent state.

Conclusion:

The investigation showed that what first looked like disposable tracing-table corruption was in fact file-level corruption of application data across several DR nodes, most likely caused by forced shutdowns on an old 2.2.5 release. The customer received a step-by-step rebuild plan, adapted as DR capacity changed, removed the first affected node, and proceeded with the remaining rebuilds. Going forward, the customer was advised to use graceful shutdowns only, keep tracing disabled, run nodetool repair on all nodes at least weekly, check disk health on affected hosts, and plan an upgrade to the final 2.2.x release and then to 3.11 or 4.x.