×
Register Here to Apply for Jobs or Post Jobs. X

CaaS Private Site Reliability Lead Engineer - Vice President

Remote / Online - Candidates ideally in
Cary, Wake County, North Carolina, 27518, USA
Listing for: Deutsche Bank AG
Full Time, Remote/Work from Home position
Listed on 2026-07-16
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 125000 - 185000 USD Yearly USD 125000.00 185000.00 YEAR
Job Description & How to Apply Below
Job Description:

J ob

Title:

CaaS Private Site Reliability Engineer Corporate

Title:

Vice President

Location:

Cary, NC

Who we are:

In short – an essential part of Deutsche Bank’s technology solution, developing applications for key business areas.

Our Technologists drive Cloud, Cyber and business technology strategy while transforming it within a robust, hands-on engineering culture. Learning is a key element of our people strategy, and we have a variety of options for you to develop professionally. Our approach to the future of work champions flexibility and is rooted in the understanding that there have been dramatic shifts in the ways we work.

Having first established a presence in the Americas in the 19th century, Deutsche Bank opened its US technology center in Cary, North Carolina in 2009. Learn more about us here .Overview As a CaaS Private Site Reliability Engineer, you will lead reliability, resilience, and operational excellence for the CaaS Private platform in US. You will bring production engineering discipline to Kubernetes, observability, automation, and incident management while helping teams improve service health and platform readiness.

You will partner across engineering, operations, and application teams to strengthen SLOs, reduce manual intervention, and ensure platform changes are measurable, supportable, and aligned to enterprise reliability standards.

What We Offer YouA diverse and inclusive environment that embraces change, innovation, and collaborationA hybrid working model, allowing for in-office / work from home flexibility, generous vacation, personal and volunteer days Employee Resource Groups support an inclusive workplace for everyone and promote community engagement

Competitive compensation packages including health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits

Educational resources, matching gift and volunteer programs

What You’ll DoLead the reliability strategy for the CaaS Private platform in US, including SLO frameworks, operational standards, and incident management maturity

Drive resilience improvements across observability, capacity planning, upgrade safety, disaster readiness, and operational automation

Own complex production management issues by leading troubleshooting, identifying root causes, and implementing preventive fixes

Define and refine service indicators, alert thresholds, dashboard standards, production readiness criteria, escalation paths, and postmortem follow-through

Develop automation and self-healing workflows that reduce manual intervention, improve recovery times, and strengthen platform supportability

Partner with cross-functional teams to ensure platform changes are measurable, supportable, and aligned with reliability objectives

How You’ll Lead Mentor junior and middle engineers while fostering a culture of blameless learning, measurable reliability, and operational excellence

Influence platform architecture and roadmap decisions using data from incidents, capacity models, operational trends, and reliability metrics

Communicate strategic insights and practical recommendations to engineering, operations, and business stakeholders to guide reliability investments

Skills You’ll Need Extensive experience with Kubernetes, Linux, distributed systems reliability, and production platform operations

Strong hands-on experience with observability, monitoring, alerting, dashboarding, incident response, and root cause analysis

Proven ability to design and implement automation, self-healing workflows, operational checks, maintenance tasks, and runbook improvements

Experience defining SLOs, service indicators, alert quality standards, production readiness practices, and escalation models

Ability to lead complex reliability improvements independently while partnering across engineering, operations, and application teams

Skills That Will Help You Excelsoft
Strong communication skills with the ability to explain technical findings to technical and non-technical stakeholders

Sound operational judgment with the ability to balance urgency, risk, and long-term platform stability during incidents

Growth mindset with a focus on…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary