Senior Site Reliability Engineer
Job in
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listed on 2026-08-22
Listing for:
CRC Group
Full Time
position Listed on 2026-08-22
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full‑time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices.
This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into Service Now and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self‑healing operations.
Location:
This role is based in Charlotte, NC.
- Define and own the enterprise monitoring and SRE observability strategy
- Serve as the subject matter expert for Dynatrace, Service Now integration, and alerting architecture
- Evaluate and recommend tooling, integration patterns, and platform direction
- Drive decisions on alerting philosophy, noise reduction, and signal quality improvement
- Architect and standardize end‑to‑end monitoring and SRE pipelines:
- Dynatrace → Service Now incident lifecycle
- Alert correlation, deduplication, and prioritization
- Integration with paging systems (Pager Duty, SMS, voice, Teams)
- Establish best practices for:
- Event ingestion and enrichment
- Incident routing and automated assignment
- Integration with CMDB and service mapping
- Lead adoption of SRE principles, including:
- SLIs, SLOs, and error budgets
- Reliability engineering practices across services
- Proactive monitoring and resilience design
- Champion a shift from reactive operations to proactive reliability engineering
- Influence application and platform teams to build observable, resilient systems by design
- Drive development of automated remediation and self‑healing capabilities
- Leverage Dynatrace workflows, Azure services, and automation frameworks to:
- Reduce manual incident handling
- Minimize unnecessary paging
- Own integration between Dynatrace and Service Now ITSM/ITOM, including:
- Incident, Event Management, and CMDB alignment
- Service mapping and dependency visibility
- Define standards for:
- Automated incident creation and resolution
- Priority assignment and routing logic
- Provide technical leadership and mentorship across SRE, platform, and application teams
- Act as a central point of coordination between engineering, cloud, and ITSM teams
- Lead workshops and working sessions to:
- Align teams on reliability practices
- Establish KPIs and drive improvement in:
- Incident response and resolution times
- Alert quality and paging effectiveness
- Monitoring coverage across critical services
- Provide leadership with clear visibility into service health and reliability trends
- 7+ years in Site Reliability Engineering, monitoring, or production engineering
- Proven experience in a technical leadership or lead engineer role
- Deep hands‑on experience with:
- Dynatrace (or equivalent observability platforms)
- Service Now ITSM / ITOM (incident, event management, CMDB)
- Demonstrated ability to:
- Design and lead enterprise monitoring/SRE architectures
- Drive platform and tooling decisions
- Integrate observability, ITSM, and paging solutions
- Experience leading SRE or observability transformation initiatives
- Strong expertise with Dynatrace–Service Now integrations
- Experience modernizing or consolidating paging/on‑call tooling
- Familiarity with:
- Azure‑based SRE tooling or AI‑assisted operations
- Infrastructure as Code (Terraform, ARM, Bicep)
- Reduction in alert noise and unnecessary paging
- Improved incident routing accuracy and MTTR
- Increased adoption of self‑healing and automated workflows
- Strong alignment between monitoring, CMDB, and service ownership
- Enterprise‑wide adoption of SRE and monitoring standards
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×