Jump to content

Mahuaa cluster issue

From TetraWiki
Revision as of 07:14, 3 December 2013 by Shashanksharma (talk | contribs)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
 Scenerio : 
     There are two nodes (mail.mahuaatv.com & mail1.mahuaatv.com) of cluster running at Mahuaa. Basically what happened was that 
     mail1.mahuaatv.com stopped working. Neither approachable via ping nor by ssh. 
 What We Did To Resue the Issue : 
 >  Shutted down both the systems. 
 >  Node 1 was brought up. 
 >  After proper startup of node 1, node 2 was brought up.
 >  Here we face an issue that node 1 was not coming up and stuck at boot process. Issue was regarding the disk issue. 
 >  Then finally after 2-3 tries, We removed lan cable & it automatically started.
 >  Both machines were up but neither if them was in network(i.e. not accessible via global ip)
 >  Then machine was brought in local n/w and it was accessed via taking teamviewer of a local machine and through that machine to server.
 >  Then following was done on server  :

On Node 2 :

 /etc/init.d/heartbeat stop
 Stopped heartbeat
 tailf /var/log/ha-log
 checked logs wether the heartbeat is properly shutting down or not
 cat /proc/drbd
 checked wether this machine is now secondary or not.
 df -h
 checked wether the /home is unmounted or not. 

now to Node 1 :

 /etc/init.d/heartbeat stop
 /etc/init.d/heartbeat start
 restart the service 
 tail -f /var/log/ha-log
 checked the logs wether the heartbeat is properly started or not. 
 cat /proc/drbd 
 check wether the server is primary or not.
 cat /etc/ha.d/haresources 
 in this file we check what ip has to be assigned to node 1 , what services to be started, which partition to be mounter and where.
 now check ip address 
 ifconfig 
  
 Check wether the /home partition is mounted or not.
 df-h

now check wether all the services are up or not which are defined in /etc/ha.d/haresources /etc/ha.d/haresources file.

example :

[root@mail ~]# cat /etc/ha.d/haresources
mail.mahuaatv.com  IPaddr::193.168.0.222/24/eth1 drbddisk::r0 Filesystem::/dev/drbd0::/home::ext3 mysqld dovecot clam_permissions clamd amavisd  postfix httpd crond
[root@mail ~]# /etc/init.d/mysqld status 
mysqld (pid 24486) is running... 
[root@mail ~]# /etc/init.d/dovecot status 
dovecot (pid  24536) is running... 
[root@mail ~]# /etc/init.d/clamd status 
clamd (pid 24617) is running... 
[root@mail ~]# /etc/init.d/amavisd status 
amavisd (pid 25009 25008 25006 25004 25003 25002 25001 25000 24999 24998 24660) is running... 
amavis-milter is stopped 
[root@mail ~]# /etc/init.d/postfix status 
master (pid 24743) is running... 
[root@mail ~]# /etc/init.d/httpd  status 
httpd (pid  24785) is running... 
[root@mail ~]# /etc/init.d/crond status 
crond (pid  24823) is running... 


now check wether the node is primary and is uptodate :

[root@mail ~]# cat /proc/drbd
version: 8.0.16 (api:86/proto:86)
GIT-hash: d30881451c988619e243d6294a899139eed1183d build by mockbuild@v20z-x86-64.home.local, 2009-08-22 13:23:56
0: cs:Connected st:Primary/Secondary ds:UpToDate/UpToDate C r---
ns:30232760 nr:8 dw:29180096 dr:2144235 al:35473 bm:787 lo:0 pe:0 ua:0 ap:0
resync: used:0/61 hits:0 misses:0 starving:0 dirty:0 changed:0
act_log: used:0/257 hits:7259549 misses:36475 starving:0 dirty:1002 changed:35473


---DONE---