Observability Architect - Seattle, Alpharetta or Cincinnati
Listed on 2026-07-31
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations
DTS is looking for Observability Architect for our Client position based in Seattle, Alpharetta or Cincinnati.
OverviewObservability & Enterprise Monitoring Architect with specialized expertise in Solar Winds platform architecture, design, and broader multi-tool observability ecosystems. Working knowledge of Open Text NNMi will be an added advantage. This role will be responsible for the end-to-end architecture, deployment, implementation, optimization, integration, and operational governance of enterprise-scale implementations of monitoring solutions (Solar Winds). Responsible for deploying platform infrastructure, establishing platform health standards, architecting automated alert workflows, designing hybrid/cloud monitoring integrations, and collaborating closely with cross-functional infrastructure and leadership teams to ensure high availability, scalability, and performance.
Roles & Responsibilities- Platform Architecture, Deployment & Lifecycle Management (Solar Winds)
- Core Module Architecture & Deployment:
Design, deploy, configure, and optimize Solar Winds modules including NPM, NCM, NTA, SAM, and the broader Orion / SWOSH (Hybrid Cloud Observability) platform ecosystem. - Deployment & Upgrade Strategy:
Lead new platform rollouts, migrations, and routine/major version updates across platform components; establish standards for platform health governance using Active Diagnostics and My Deployment health checks. - Polling Infrastructure Deployment:
Architect, deploy, scale, and load-balance Additional Polling Engines (APEs) to ensure optimal performance, redundancy, and capacity across enterprise environments. - Database & Storage Strategy:
Oversee architectural strategy for the underlying MS SQL Database, ensuring high availability, performance tuning, and robust configuration and database backup governance. - Network & Device Monitoring
- Discovery & Asset Onboarding:
Execute network discoveries, deploy automated node onboarding/offboarding frameworks, assign Universal Device Pollers (UnDP), and maintain custom attribute taxonomies and group hierarchies. - Configuration Governance (NCM):
Design and implement NCM command templates, establish policies for automated daily startup/running config backups, config archiving, and remediation frameworks for compliance/transfer failures. - Topology & Visualization:
Build dynamic, accurate network topology frameworks using Network Atlas and modern visual canvases aligned with enterprise requirements. - Alert Architecture, Dashboarding & ITSM Integration
- Signal & Alert Optimization:
Design, implement, and tune custom Alert Triggers, Actions, and Threshold frameworks to eliminate alert noise and establish high-signal, actionable alerting. - ITSM & Workflow Deployment:
Deploy bi-directional ITSM/ticketing integrations to enable automated ticket creation, enrichment, routing, and lifecycle tracking. - Reporting & Visibility Frameworks:
Build enterprise operational and executive Dashboards, Views, and Reports tailored to multi-level stakeholder requirements. - Incident & Deployment Support:
Lead technical reviews for complex operational anomalies, troubleshoot systemic telemetry or deployment issues, and collaborate with domain teams on root cause analysis (RCA). - AIOps & Next-Gen Operations
- AIOps Implementation:
Define, deploy, and leverage AIOps, machine learning, and pattern-recognition capabilities to identify baseline anomalies, reduce event noise, and drive predictive incident management. - Automated Remediation Architecture:
Collaborate with cross-functional teams to integrate AI-driven event correlation models and deploy automated self-healing remediation workflows into the central monitoring platform. - Integration, Vendor Coordination
- Integration, Vendor Coordination
- API & Integration Deployment:
Implement REST API and webhook integration models across enterprise applications, tools, and platforms as per business requirements. - Troubleshooting & Diagnostics
- Advanced Escalation:
Perform deep-dive troubleshooting and root-cause analysis for complex, platform-level performance degradations, engine polling deadlocks, and monitoring agent corruptions. - Telemetry Diagnostics:
Utilize Active Diagnostics and system telemetry data to investigate and resolve complex network configuration transfer failures, polling sync latency, and data ingestion bottlenecks.
- Multi-tool Architecture & Deployment Expertise:
Deep architectural and hands-on implementation knowledge of enterprise monitoring tools (Solar Winds, Open Text, Splunk, etc.) at global scale. - Protocol & Telemetry Mastery:
Advanced understanding of SNMP (v2c/v3), WMI, WinRM, Syslog, Net Flow/sFlow, and core Observability pillars (Metrics, Logs, Traces). - Automation & API Design:
Intermediate skills in Power Shell/Python, REST APIs, and building API-driven automation for enterprise monitoring workflows. - AIOps & Intelligent Automation:
Strong grasp of AIOps concepts, machine learning algorithms for anomaly detection, automated event correlation, and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).