Browse Definitions :
Definition

split brain syndrome

Contributor(s): Arun prasath S

Split brain syndrome, in a clustering context, is a state in which a cluster of nodes gets divided (or partitioned) into smaller clusters of equal numbers of nodes, each of which believes it is the only active cluster. 

Believing the other clusters are dead, each cluster may simultaneously access the same application data or disks, which can lead to data corruption. A split brain situation is created during cluster reformation. When one or more node fails in a cluster, the cluster reforms itself with the available nodes. During this reformation, instead of forming a single cluster, multiple fragments of  the cluster with an equal number of nodes may be formed. Each cluster fragment assumes that it is the only active cluster -- and that other clusters are dead -- and starts accessing the data or disk. Since more than one cluster is accessing the disk, the data gets corrupted.

Here's how it works in more detail:

  • Let's say there are 5 nodes A,B,C,D and E which form a cluster, X.
  • Now a node (say E) fails.
  • Cluster reformation takes place. Actually, the remaining nodes A,B,C and D should form cluster X.
  • But split brain situation may occur which leads to formation of two clusters X1 (containing A and B) and X2 (containing C and D).
  • Both X1 and X2 clusters think that they are the only active cluster. Both clusters start accessing the data or disk, leading to data corruption.

High availability clusters are all vulnerable to split brain syndrome and should use some mechanism to avoid it. Clustering tools, such as Pacemaker, HP ServiceGuard, CMAN and LinuxHA, generally include such mechanisms.

Common methods of addressing split brain syndrome include:

 

This was last updated in March 2014

Continue Reading About split brain syndrome

Start the conversation

Send me notifications when other members comment.

Please create a username to comment.

SearchCompliance

  • risk assessment

    Risk assessment is the identification of hazards that could negatively impact an organization's ability to conduct business.

  • PCI DSS (Payment Card Industry Data Security Standard)

    The Payment Card Industry Data Security Standard (PCI DSS) is a widely accepted set of policies and procedures intended to ...

  • risk management

    Risk management is the process of identifying, assessing and controlling threats to an organization's capital and earnings.

SearchSecurity

SearchHealthIT

SearchDisasterRecovery

  • call tree

    A call tree is a layered hierarchical communication model that is used to notify specific individuals of an event and coordinate ...

  • Disaster Recovery as a Service (DRaaS)

    Disaster recovery as a service (DRaaS) is the replication and hosting of physical or virtual servers by a third party to provide ...

  • cloud disaster recovery (cloud DR)

    Cloud disaster recovery (cloud DR) is a combination of strategies and services intended to back up data, applications and other ...

SearchStorage

Close