Jump to content

RCA REPORT FOR RAMNATHPURA - NIIT Tobacco Cluster Failure

From TetraWiki
Revision as of 13:44, 2 February 2013 by Biswajit (talk | contribs)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)


Root Cause Analysis (RCA)


#


Location Ramnathpuram , Karnataka , Plaform No 63
#


Incident ticket Number 2013020110000595
#


Priority 1
#


Date & Time of Incident Reported 01-Feb-2013 11:06 hrs
#


RCA Date 02-Feb-2012
#


Description of the Incident and how the incident was detected Platform services was not available

Customer reported



#


Impact of the Incident Platform 63 HHT services was not available
#


Investigation Findings
    • What happened
    • How did it happen
    • Why it happened
    • Recommendations


11:06 AM – Phone call .

11.15 to 13 PM – Screen shot send to Rajneesh ( not to Service help desk ) . Rajneesh has guided the local team ( Raj kumar ) to make cluster up .

13:00 – 16:00 – Tetra remote support desk got involved and tried to diagnose the issue. Due to lack of basic Linux knowledge at the local end and no team-viewer access , major delay in understanding the issue by Tetra Team .

16:00 – Decision was taken to manual start the services on Node 2

18:30 PM – Node 2 was on-line (with out cluster ) and Platform services was Up with Single Node

19:00 – 04:00 – Efforts to make the cluster up . Senior Engineer , Bhushan , reached site. Provided logs . Logs were analyzed .

8:00 AM – Production time started and one node was on production . Tetra team was still analyzing and cause and fix .

16:30 PM – production hours were over .

16:45 PM – Cluster was tested UP and running

Analyzing the logs , the following conclusion was inferred

  • The Cause of issue was repeated power failures . This has made data sync between the cluster node , inconsistent
  • Logs also revealed that process inconsistency on process DRBD.
  • It was also discovered that drbd process was not on for both the nodes at the boot time .
  • We had forcefully made the data consistent .
  • Switching on the DRBD process on boot time made the cluster run .
  • Logs are attached


#


Recommended measures for prevention of similar incidents
  • The recommendation should consider the following:
  1. Any process improvement
  2. Any technical measures
  3. Any training/awareness


After analysis we recommend that
  • It is 100% mandatory to have proper UPS system . The power should be available 100% without fail to the Cluster
  • Team viewer access
  • At least 3-6 month experience Linux L0 resources , major delay has happen due to no basic knowledge
  • Remote training to L0/L1 local resource to have fast turnaround
  • Monthly health report generation


#


Follow up actions on prevention of similar incidents * SOS UPS availability
  • Training to L0 / L1 Resource .