Monitoring Engineer
Listed on 2026-07-27
-
IT/Tech
Cloud Computing: Infrastructure & Operations, IT Support, Systems Administrator, Cybersecurity
Senior Observability & Enterprise Monitoring Engineer
Alpharetta, Georgia/Cincinnati, Ohio/Seattle, WA (Local candidates only)
Overview:
Observability & Enterprise Monitoring Engineer with specialized expertise in Solar Winds platform administration and broader multi-tool observability ecosystems. Working knowledge of Open Text NNMi will be an added advantage. This role will be responsible for the end-to-end administration, optimization, integration, and operational maintenance of enterprise-scale implementation of monitoring solutions (Solar Winds). Responsible for ensuring platform health, automating alert workflows, managing hybrid/cloud monitoring integrations, and collaborating closely with cross-functional infrastructure teams to maintain high availability and performance.
1.Platform Administration & Lifecycle Management (Solar Winds)
Core Module Management:
Administer and optimize Solar Winds modules, including NPM, NCM, NTA, SAM, and the broader Orion / SWOSH (Hybrid Cloud Observability) platform ecosystem.
Upgrades & Maintenance:
Perform routine and major version updates across platform components; monitor platform health using Active Diagnostics and My Deployment health checks.
Polling Infrastructure:
Manage, scale, and load-balance Additional Polling Engines (APEs) to ensure optimal performance across enterprise environments.
Database & Backup Operations:
Perform operational tasks on the underlying MS SQL database and manage, schedule, and verify configuration and database backup jobs.
Discovery & Asset Management:
Execute network discoveries; manage node onboarding/offboarding; assign Universal Device Pollers (UnDP); and maintain custom attributes and group hierarchies.
Configuration Management (NCM):
Build and maintain NCM command templates, automate daily startup/running config backups, archive config files, and remediate compliance/transfer failures.
Topology & Visualization:
Create and maintain dynamic, accurate network topology maps using Network Atlas and modern visual canvases based on operational requirements.
Signal Optimization:
Design, tune, and maintain custom Alert Triggers, Actions, and Thresholds to eliminate alert noise and drive actionable alerting.
Ticketing & Automation:
Configure bi-directional ITSM/ticketing integrations to enable automatic ticket creation, routing, and lifecycle tracking.
Reporting & Visibility:
Build custom operational and executive Dashboards, Views, and Reports tailored to stakeholder requirements.
Incident Support:
Monitor alert channels for operational anomalies, troubleshoot lingering telemetry issues, and collaborate with domain teams to drive root cause resolution.
Leverage AIOps, machine learning, and pattern-recognition capabilities to identify baseline anomalies, reduce event noise, and drive predictive incident management.
Collaborate with cross-functional teams to integrate AI-driven event correlation models and automated remediation workflows into the central monitoring platform.
Manage relationships and support escalations with platform vendors.
Work on REST API integrations across applications/tools as per requirements.
6. Operational Troubleshooting & DiagnosticsPerform deep-dive troubleshooting and root-cause analysis for platform-level performance degradations, engine polling failures, and monitoring agent corruptions.
Utilize Active Diagnostics and system telemetry to investigate and resolve complex network configuration transfer failures, polling sync latency, and data ingestion issues.
Required Skills:Protocol & Telemetry Knowledge:
In-depth understanding of SNMP (v2c/v3), WMI, WinRM, Syslog, Net Flow/sFlow, and observability (Metrics, Logs, Traces).
Automation & API Integration:Good to have skills in Power Shell/Python and API-driven automation for monitoring workflows.
AIOps & Intelligent Automation:Basic understanding of AIOps concepts, machine learning algorithms for anomaly detection, automated event correlation, and predictive analytics within modern observability frameworks.
Cloud & Hybrid Observability:Hands-on experience extending platform monitoring into AWS, Azure, or GCP environments.
Infrastructure Knowledge- System Administration:
Intermediate knowledge of Windows and Linux administration. - Database:
Understanding of SQL/database concepts and standard query execution. - Networking:
Good understanding of networking concepts including TCP/IP, DNS, DHCP, routing, and switching. - ITSM:
Experience in ITSM processes and operational support.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).