Head of Enterprise Site Reliability Engineering; SRE; Hybrid
Listed on 2026-08-31
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
AVP / Head Of Enterprise Site Reliability Engineering (SRE)
Genesis
10 is currently seeking a AVP / Head of Enterprise Site Reliability Engineering (SRE) with our consumer finance lender firm client in their Pittsburgh, PA location. This is a right to hire position. An experienced Site Reliability Engineering leader is needed to establish, lead, and scale a formal enterprise SRE capability. This is a contract-to-hire opportunity intended to convert to a full-time AVP-level position based on performance, organizational approval, and mutually agreed-upon terms.
This leader will define the enterprise SRE strategy, operating model, governance framework, roadmap, and success measures while building and managing a centralized team focused on reliability, resilience, operational maturity, and engineering velocity across on-premises, hybrid, and cloud environments. The role will drive the adoption of Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, observability standards, automation frameworks, incident management practices, and production-readiness expectations.
The successful candidate will establish reliability as a measurable business and engineering outcome by creating repeatable patterns, paved-road standards, and operating practices that reduce toil and improve service performance. This is a strategic, hands-on leadership position requiring strong people leadership, executive communication, program development, and cross-functional influence. The leader will partner with senior stakeholders across Technology, Product, Architecture, Security, Risk, Compliance, and Operations to make reliability a core product and platform capability.
Responsibilities:
- Establish the enterprise SRE strategy, vision, roadmap, operating model, and governance framework in alignment with business priorities, technology modernization, cloud adoption, and operational resilience goals.
- Build, manage, and develop a centralized SRE team, including defining roles, creating staffing plans, establishing performance expectations, coaching team members, supporting career development, and planning for succession.
- Define the SRE engagement model, including service tiers, onboarding criteria, production-readiness standards, intake and prioritization processes, and partnership expectations for product, platform, infrastructure, and application teams.
- Define and operationalize enterprise reliability measures, including SLIs, SLOs, error budgets, availability targets, toil-reduction goals, incident metrics, change failure rate, mean time to detect, and mean time to restore.
- Establish error-budget policies and executive-level decision frameworks that guide tradeoffs among delivery velocity, operational risk, and service stability.
- Develop and govern enterprise reliability standards and reusable assets, including golden-signal frameworks, observability patterns, alerting standards, runbook templates, SLO dashboards, incident playbooks, and production-readiness checklists.
- Advance incident management maturity through effective high-severity response, executive communications, blameless post-incident reviews, root-cause analysis, corrective-action tracking, and measurable reliability improvements.
- Oversee operational readiness, capacity planning, disaster-recovery validation, resilience testing, service-health reviews, and reliability risk management for supported platforms and critical business services.
- Drive automation and observability strategies that reduce operational toil, improve visibility, accelerate recovery, and enable scalable support models across hybrid and cloud platforms.
- Identify, sponsor, and govern AI-enabled reliability capabilities, including alert-noise reduction, incident summarization, event correlation, runbook assistance, predictive operations, and approved automated remediation.
- Ensure AI-enabled operational capabilities are implemented with appropriate privacy, security, risk, compliance, auditability, and human-oversight controls.
- Partner with Security, Risk, Compliance, Audit, Architecture, Product, Development, Infrastructure, Cloud Operations, and Technology Operations leaders to embed reliability requirements throughout the software development lifecycle and production support model.
- Communicate reliability posture, risks, progress, business impact, and investment requirements to executive stakeholders using clear metrics and actionable recommendations.
- Promote shared ownership, engineering excellence, accountability, continuous improvement, and blameless learning across technology teams.
- Remain current on SRE, observability, platform engineering, AIOps, resilience engineering, and cloud reliability practices, applying relevant approaches to improve business and operational outcomes.
- Provide leadership during major incidents and critical operational events, including escalation management, cross-functional coordination, stakeholder communications, and executive updates.
Requirements:
- Progressive…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).