Jump to content
Main menu
Main menu
move to sidebar
hide
Navigation
Main page
Recent changes
Random page
Help about MediaWiki
TetraWiki
Search
Search
Appearance
Create account
Log in
Personal tools
Create account
Log in
Pages for logged out editors
learn more
Contributions
Talk
Editing
Spectra Nagios Project Completion Report
Page
Discussion
English
Read
Edit
View history
Tools
Tools
move to sidebar
hide
Actions
Read
Edit
View history
General
What links here
Related changes
Special pages
Page information
Appearance
move to sidebar
hide
Warning:
You are not logged in. Your IP address will be publicly visible if you make any edits. If you
log in
or
create an account
, your edits will be attributed to your username, along with other benefits.
Anti-spam check. Do
not
fill this in!
''Auto-generated from the uploaded PDF [[:File:Spectra_Nagios_Project_Completion_Report.pdf|Spectra_Nagios_Project_Completion_Report.pdf]]. This is an extracted-text rendering for searchability; see the original PDF for exact formatting, diagrams, tables, and images.'' <pre> <nowiki> Project Completion Report Project Title: Nagios XI Upgrade and HA Implementation Project Duration: March 2025 β June 2025 Prepared For: spectra.co Prepared By: Tetra Information Services LTD. Date: 03 July 2025 Table of Contents 1. Executive Summary 2. Project Objectives 3. Scope of Work 4. Implementation Timeline 5. Detailed Activities and Outcomes o Module 1: Design HA Architecture o Module 2: OS and Application Upgrade o Module 3: HA Configuration for Main o Module 4: HA Configuration for Probe o Module 5: Database Issue Resolution o Module 6: Main-Probe Synchronization o Module 7: API & Integration o Module 8: Hardware and Software Planning 6. Key Challenges & Resolutions 7. Recommendations 8. Conclusion 9. Appendices 1. Executive Summary This report outlines the successful completion of the Nagios XI Upgrade and High Availability (HA) Implementation project carried out for spectra.co by Tetra Information Services LTD between March and June 2025. The goal was to revamp the monitoring system by upgrading the outdated CentOS-based setup, introducing redundancy through HA configurations, and improving reliability, database performance, and scalability. This project aligns the system for anticipated growth over the next 3β5 years and provides a resilient infrastructure to support 15,000+ hosts and over 150,000 monitored services. All modules planned under this initiative have been executed and validated in collaboration with the client. Deliverables include new HA architecture, upgraded OS/Nagios versions, optimized databases, real-time configuration sync, and hardware recommendations for future scaling. 2. Project Objectives 2.1 Core Objectives ο· To implement a robust and fault-tolerant High Availability (HA) configuration for both Main and Probe monitoring nodes. ο· To upgrade the Operating System from CentOS 7 (EOL) to Ubuntu 20.04 LTS. ο· To upgrade Nagios XI from version 5.8.9 to the latest stable release, 2024R1.4.4. ο· To address database performance issues, backup failures, and table bloating. ο· To ensure real-time synchronization between Main and Probe environments. ο· To stabilize integration with API and Service Request (SR) automation tools. 2.2 Key Deliverables ο· HA design documentation finalized with the client ο· Clean OS installations and configuration ο· Upgraded Nagios XI instances ο· Master-Master MySQL Replication and DRBD setup ο· Automated database cleanup configurations ο· Resolved main-probe sync lags ο· Reviewed and stabilized API/SR behavior ο· Capacity roadmap to scale up to 150K services 3. Scope of Work The project scope included infrastructure redesign, HA implementation, and full-stack upgrades with minimal downtime. The following major components were included: ο· Planning and designing of Main/Probe HA infrastructure ο· Execution of OS and Nagios XI upgrades on both nodes ο· Re-establishing DRBD and MySQL replication clusters ο· Resolving historical database errors ο· Ensuring production-ready probe-main synchronization ο· Addressing post-upgrade API/SR issues ο· Providing a future hardware expansion strategy 4. Implementation Timeline Instead of a table, hereβs the timeline as a step-wise breakdown with date markers: ο· 19β25 March 2025 β HA architecture finalized with Spectra ο· 26 March β 07 May 2025 β SOC (Secondary) Node OS and Nagios installation completed ο· 28 March β 13 May 2025 β Configuration and database migration to new node ο· 16β19 May 2025 β Switchover to new SOC node completed ο· 20β23 May 2025 β Primary Node OS/Nagios rebuilt and upgraded ο· 24 May β 10 June 2025 β Main and Probe HA clusters established ο· 10β25 June 2025 β Final testing, sync tuning, and hardware roadmap delivered Each activity followed strict checklists, supported by configuration screenshots and validations reviewed by the client. 5. Detailed Activities and Outcomes Module 1: HA Architecture Design ο· Multiple HA models were proposed, tested, and validated. ο· Final decision involved using MySQL master-master replication and DRBD for file-level sync. ο· No major new hardware purchases were required at this stage. ο· Documentation of cluster failover scenarios was shared with Spectra. Module 2: OS and Application Upgrade 2.1 HA Break and Initial Preparation ο· HA was temporarily disabled. ο· Services temporarily shifted to the Main node. 2.2 Secondary Node Installation ο· Ubuntu 22.04 LTS was first used, but changed to 20.04 LTS as per client feedback. ο· Nagios XI 5.7 installed successfully, preparing for upgrade. 2.3 Configuration Migration ο· Nagios configurations, DBs, custom scripts were moved from the old system. ο· Data validation, backup recovery, and plugin compatibility tests were performed. 2.4 SOC Go-Live ο· Final production cutover took place on 19 May 2025. ο· Monitoring services resumed without loss of data or downtime. 2.5 Primary Node Upgrade ο· Rebuilt Primary Node from scratch with Ubuntu 20.04 LTS. ο· Upgraded Nagios XI to 2024R1.4.4. ο· Configured DRBD and MySQL replication to secondary node. Module 3: HA Configuration β Main ο· Cluster fencing, shared IP configuration, and failover scripting implemented. ο· DRBD replication verified for logs, configs, and backups. ο· MySQL master-master pair synchronized and validated. ο· All services continued without interruption during switchover simulations. Module 4: HA Configuration β Probe ο· Probe node was duplicated and configured with identical packages. ο· MySQL cluster and DRBD replication implemented as in Main node. ο· Configured to handle real-time alerts and monitoring data from remote devices. Module 5: Database Issue Resolution ο· Backups were previously failing due to large .ibd files and table locks. ο· Old tables (over 5GB each) were split and purged with automation. ο· Enabled automated cleanup scripts in Nagios XI. ο· Post-cleanup system is stable, with backup jobs running nightly. Module 6: MainβProbe Synchronization ο· Prior to the upgrade, there were 15β20 minute lags in alarm sync. ο· Post-upgrade, the sync is real-time, with instant propagation of changes. ο· This led to faster deployment, reduced troubleshooting, and higher monitoring accuracy. Module 7: API & Integration ο· SR automation previously created duplicate or missing records. ο· Integration scripts were revalidated and corrected with vendor support. ο· Delay in SR acknowledgment reduced from 5β10 mins to <2 mins on average. ο· Final integration success rate is now above 97%. Module 8: Hardware and Software Planning ο· Final architecture allows for growth up to 150K services. ο· Main HA servers: / = 600GB, /data = 5.2TB, RAM = 128GB, vCPU = 12 ο· Probe nodes: / = 500GB, /data = 500GB, RAM = 32GB, vCPU = 4β6 ο· Hardware scaling roadmap provided and accepted by Spectra. 6. Key Challenges & Resolutions ο· OS version conflict resolved by switching from 22.04 LTS to 20.04 LTS. ο· DRBD sync stalls fixed by tuning resource parameters. ο· Partition size mismatches addressed via disk resizing. ο· Backup failures resolved through database optimization. 7. Recommendations ο· Continue monitoring synchronization and replication status weekly. ο· Maintain quarterly health checks on MySQL size and HA logs. ο· Keep probe configuration updated for future monitoring growth. ο· Begin planning for cloud-based hybrid Nagios XI extension in 2026. 8. Conclusion The project has been executed successfully and fulfills all design, functionality, and availability expectations. The Nagios XI infrastructure is now fault-tolerant, stable, and scalable for years to come. Final Handover Date: 25 June 2025 Project Closure Status: β Fully Completed 9. Appendices 9.1 Reference Document List The following documents were created, reviewed, and shared with spectra.co throughout the project lifecycle. Each document supports the execution and validation of key modules in the HA architecture, upgrade strategy, configuration, and optimization. 1. Proposed Monitoring Architecture β Spectra Describes the final HA design including main and probe node layout, resource planning, and IP failover strategy. 2. Solution Document β Nagios HA Comparison with Architecture Diagram Provides comparative analysis of different HA deployment methods and justifies selected approach for Spectra. 3. PCS Cluster Failover β Process Step-by-step documentation outlining cluster fencing, VIP migration, and automatic failover process validation. 4. Spectra_SOC Nagios Optimization β Project Status (Final Sheet) Status tracking sheet showing all completed activities, timelines, dependencies, and mitigation actions. 5. Nagios XI Cluster Setup Checklist Comprehensive deployment checklist used by engineers to ensure no step was missed during HA and sync configuration. </nowiki> </pre>
Summary:
Please note that all contributions to TetraWiki may be edited, altered, or removed by other contributors. If you do not want your writing to be edited mercilessly, then do not submit it here.
You are also promising us that you wrote this yourself, or copied it from a public domain or similar free resource (see
TetraWiki:Copyrights
for details).
Do not submit copyrighted work without permission!
Cancel
Editing help
(opens in new window)