×
Register Here to Apply for Jobs or Post Jobs. X

CaaS Private Site Reliability Engineer - Assistant Vice President

Remote / Online - Candidates ideally in
Cary, Wake County, North Carolina, 27518, USA
Listing for: Deutsche Bank AG
Full Time, Remote/Work from Home position
Listed on 2026-07-16
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 100000 - 153000 USD Yearly USD 100000.00 153000.00 YEAR
Job Description & How to Apply Below
Job Description:

J ob Title CaaS Private Site Reliability Engineer Corporate Title Assistant Vice President Location Cary, NC

Who we are:

In short – an essential part of Deutsche Bank’s technology solution, developing applications for key business areas.

Our Technologists drive Cloud, Cyber and business technology strategy while transforming it within a robust, hands-on engineering culture. Learning is a key element of our people strategy, and we have a variety of options for you to develop professionally. Our approach to the future of work champions flexibility and is rooted in the understanding that there have been dramatic shifts in the ways we work.

Having first established a presence in the Americas in the 19th century, Deutsche Bank opened its US technology center in Cary, North Carolina in 2009. Learn more about us here .Overview As a Site Reliability Engineer on the CaaS Private platform team, you will help operate and improve an on-prem, multi-tenant Kubernetes platform running on bare metal. You will strengthen the reliability, observability, scalability, and operational excellence of a platform that supports critical, low-latency, and regulated workloads.

You will partner closely with platform, network, security, and application teams to define service level objectives, improve resilience, reduce operational toil, and build automation that allows the platform to run safely n us here, and you will turn operational challenges into measurable engineering improvements that application teams can rely on every day.

What We Offer YouA diverse and inclusive environment that embraces change, innovation, and collaborationA hybrid working model, allowing for in-office / work from home flexibility, generous vacation, personal and volunteer days Employee Resource Groups support an inclusive workplace for everyone and promote community engagement

Competitive compensation packages including health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits

Educational resources, matching gift and volunteer programs

What You’ll DoDefine, implement, and continuously improve SLI/SLOs, alerting standards, and error budgets for the CaaS Private platform and critical services

Build and maintain observability across metrics, logs, alerts, and dashboards to provide clear insight into platform health, saturation, latency, and failure modes

Lead or coordinate incident response for platform-impacting events, ensuring timely mitigation, clear communication, blameless postmortems, and durable follow-up actions

Automate repetitive operational tasks and remediation workflows to reduce toil, improve platform consistency, and accelerate recovery time Improve reliability, upgrade safety, and operational readiness for Kubernetes clusters, ingress paths, service mesh components, node services, and critical platform dependencies

Partner with platform, network, security, and application teams on capacity planning, release readiness, troubleshooting, operational documentation, and adoption of best practices

Skills You’ll Need Proven experience in Site Reliability Engineering, Production Engineering, Dev Ops, or a closely related infrastructure role Hands-on Kubernetes expertise, including operating clusters on bare metal or private cloud environments and supporting platform services at scale

Strong Linux system administration capability and infrastructure-level scripting experience using Python, Ansible, and Bash Practical knowledge of observability stacks and telemetry pipelines, including Prometheus, Grafana, Splunk, metrics, logging, alerting, dashboards, and Open Telemetry-style concepts

Strong understanding of incident management, root cause analysis, operational readiness, computer networking, virtualization, containerization, and distributed systems behavior under failure

Skills That Will Help You Excel

Experience with Istio / Envoy, service mesh observability, traffic management, OPA Gatekeeper, admission controls, or policy-driven operational guardrails

Familiarity supporting stateful services such as PostgreSQL, Kafka, MongoDB, or comparable platform dependencies

Practical knowledge of…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary