Site Reliability Engineer
Listed on 2026-08-05
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / Service Now)
We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices. This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into Service Now and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self-healing operations.
Key Responsibilities
- Strategic Leadership & Decision-Making
- Define and own the enterprise monitoring and SRE observability strategy
- Serve as the subject matter expert for Dynatrace, Service Now integration, and alerting architecture
- Evaluate and recommend tooling, integration patterns, and platform direction
- Drive decisions on alerting philosophy, noise reduction, and signal quality improvement
- Platform Ownership & Architecture
- Architect and standardize end-to-end monitoring and SRE pipelines:
- Dynatrace → Service Now incident lifecycle
- Alert correlation, deduplication, and prioritization
- Integration with paging systems (Pager Duty, SMS, voice, Teams)
- Establish best practices for:
- Event ingestion and enrichment
- Incident routing and automated assignment
- Integration with CMDB and service mapping
- Site Reliability Engineering (SRE) Leadership
- Lead adoption of SRE principles, including:
- SLIs, SLOs, and error budgets
- Reliability engineering practices across services
- Proactive monitoring and resilience design
- Champion a shift from reactive operations to proactive reliability engineering
- Influence application and platform teams to build observable, resilient systems by design
- Automation & Self-Healing Enablement
- Drive development of automated remediation and self-healing capabilities
- Leverage Dynatrace workflows, Azure services, and automation frameworks to:
- Reduce manual incident handling
- Eliminate repeatable operational tasks
- Minimize unnecessary paging
- Service Now & Observability Integration Leadership
- Own integration between Dynatrace and Service Now ITSM/ITOM, including:
- Incident, Event Management, and CMDB alignment
- Service mapping and dependency visibility
- Governance for application/service tagging
- Define standards for:
- Automated incident creation and resolution
- Priority assignment and routing logic
- Monitoring-to-ITSM data synchronization
- Team Leadership & Cross-Functional Influence
- Provide technical leadership and mentorship across SRE, platform, and application teams
- Act as a central point of coordination between engineering, cloud, and ITSM teams
- Lead workshops and working sessions to:
- Drive monitoring standardization
- Align teams on reliability practices
- Influence upstream architectural decisions
- Operational Excellence
- Establish KPIs and drive improvement in:
- Incident response and resolution times
- Alert quality and paging effectiveness
- Monitoring coverage across critical services
- Provide leadership with clear visibility into service health and reliability trends
Required Qualifications
- 7+ years in Site Reliability Engineering, monitoring, or production engineering
- Proven experience in a technical leadership or lead engineer role
- Deep hands-on experience with:
- Dynatrace (or equivalent observability platforms)
- Microsoft Azure (IaaS, PaaS, networking, identity)
- Service Now ITSM / ITOM (incident, event management, CMDB)
Preferred Qualifications
- Experience leading SRE or observability transformation initiatives
- Strong expertise with Dynatrace–Service Now integrations
- Experience modernizing or consolidating paging/on-call tooling
- Familiarity with:
- Azure-based SRE tooling or AI-assisted operations
- Automation frameworks (Git Hub Actions, Runbooks, etc.)
- Infrastructure as Code (Terraform, ARM, Bicep)
Success Metrics
- Reduction in alert noise and unnecessary paging
- Improved incident routing accuracy and MTTR
- Increased adoption of self-healing and automated workflows
- Strong alignment between monitoring, CMDB, and service ownership
- Enterprise-wide adoption of SRE and monitoring…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).