Expert Reliability Engineer
Listed on 2026-09-04
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, Cybersecurity, SRE/Site Reliability
We help the world run better
At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong.
What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.
******* Due to the potentially classified nature of our work, your willingness is required to subject yourself to a governmental security clearance process ********
YOUR FUTURE ROLEWe are looking for an Expert Reliability Engineer (SRE) within the Shared Management Services (SMS) group in the Technology and Engineering unit of SAP Sovereign Cloud organization.
In this role, you will join the Technology and Engineering team as a Site Reliability Engineer focused on securing and scaling the foundational platform that underpins SAP Sovereign Cloud. You will work alongside a globally distributed team of highly motivated engineers responsible for the design, development, deployment, and lifecycle management of the Sovereign Cloud Shared Management Services (SMS) platform, with a mandate that spans both operational reliability and platform security.
As an expert Site Reliability Engineer, you will shoulder a shared ownership of the reliability, security, and operational excellence of production and non-production environments within Sovereign Cloud. It is expected that you will have extensive experience in all areas of a critical application administration stack spanning:
- source control (git)
- CI/CD platforms
- identity and access management
- secrets management
- container orchestration
- network security infrastructure
- full-stack observability tooling across multi-cloud environments.
- AI-assisted engineering workflows
You will lead an experienced team of globally distributed engineers in setting standards and best practices across responsibility areas including:
- Code reviews
- Agentic coding (AI) security best practices
- Agentic coding workflows and skill development
- AI-assisted incident response
- Unit test coverage
- Functional test coverage
- etc
You will treat security as a first-class reliability concern: hardening identity and access management, secrets management, and supply chain integrity are as central to this role as uptime and incident response.
You will identify and close mission-critical capability gaps, define disciplined and standardized operational processes, and help the team navigate trade-offs across deployment plans, infrastructure investments, and day-to-day operational decisions. You will own and continuously improve backup and disaster recovery drills, ensuring the platform is failure-ready at global scale.
You will work closely with Sovereign Cloud operations and engineering teams, regional counterparts, and hyperscaler provider partners.
WHAT YOU BRING- Ability to manage ambiguities while being innovative and collaborative
- Extensive technology skills and the willingness to learn new topics quickly
- Problem-solving, presentation, communication, and interpersonal skills
- Ability to think strategically, delivering projects and work cross-organizationally
- Knowledge of SAP and the SAP solution portfolio
- Cultural awareness, intercultural competencies, and the ability to influence without formal authority
- Ability to build trusted relationships with key stakeholders
- Persistence, self-motivation, and willingness to work under pressure
- Proven ability to work in cross-functional teams
- Ability to lead and mentor junior engineers in setting and maintaining Dev Ops and SRE best practices
- English (fluent)
- 10 years of experience in Dev Ops and/or SRE engineering
- 7 years of experience with a successful track record of leading engineering projects and cross-functional program teams
- Deep mastery of SRE principles as they apply to a globally distributed, mission-critical platforms
- Advanced Experience in architecture, engineering, and deployment of modern monitoring tooling such as Grafana, Promethius
- Background in security or…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).