Site Reliability Engineer, Observability
Listed on 2026-09-20
-
IT/Tech
Cybersecurity, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
About Ripple
At Ripple, we’re building a world where value moves like information does today. It’s big, it’s bold, and we’re already doing it. Through our crypto solutions for financial institutions, businesses, governments and developers, we are improving the global financial system and creating greater economic fairness and opportunity for more people, in more places around the world.
Ripple Treasury, now a Ripple solution, was acquired by Ripple in 2025, marking a significant expansion into the multi-trillion-dollar corporate finance arena. Ripple Treasury has more than 40 years of experience supporting some of the world’s largest and most sophisticated companies. Integrating its treasury command center into Ripple’s technology stack gives corporates the ability to move, manage and optimize liquidity in real-time, across traditional and digital assets, under one expanded umbrella.
Join us to build the future of corporate treasury and the infrastructure that powers the Internet of Value.
PositionSenior Site Reliability Engineer – part of Ripple’s Technical Operations team. This role focuses on observability, releasability, security, and Dev Sec Ops practices to keep Ripple’s products highly available, performant, and resilient at scale.
Key Responsibilities- Embed with stream-aligned teams in time-boxed engagements (6–12 weeks) to build capabilities across observability, CI/CD, and security domains.
- Coach teams on instrumenting code with logs, metrics, and traces; creating dashboards, alerts, and defining SLOs/SLIs and error budgets to improve production troubleshooting.
- Guide teams in designing and optimizing CI/CD pipelines using Azure Dev Ops, Git Hub Actions, and Octopus Deploy, including trunk-based development and progressive delivery strategies.
- Teach teams to integrate security scanning (SAST, DAST, SCA) into pipelines, perform threat modeling, implement secrets management, and remediate vulnerabilities.
- Identify and remove operational bottlenecks across deployment, monitoring, and security practices to reduce lead time, MTTR, and vulnerability remediation time.
- Develop reusable templates, documentation, and golden path examples for pipelines, dashboards, and security controls that teams can adopt independently.
- Measure and track team progress on DORA metrics, observability maturity, and security posture, demonstrating measurable improvement over each engagement.
- Collaborate with the Subsystems Platform Team to translate recurring team needs into scalable, self-service platform capabilities.
- Facilitate knowledge sharing through workshops, pairing sessions, and documentation that builds lasting team competence across all three domains.
- 5+ years of experience in Site Reliability Engineering, Dev Ops, or Platform Engineering with hands‑on exposure across multiple disciplines.
- Strong expertise in application observability tools (New Relic, Datadog, or similar), including structured logging, distributed tracing, and dashboard/alert creation.
- Proven experience designing and implementing CI/CD pipelines using Azure Dev Ops, Git Hub Actions, and Octopus Deploy, with deep knowledge of deployment automation.
- Solid understanding of application security practices including SAST, DAST, SCA tools, OWASP Top 10, and Dev Sec Ops integration into CI/CD pipelines.
- Strong experience with Infrastructure as Code (IaC) using Terraform or similar tools, and knowledge of trunk-based development and version control best practices.
- Understanding of SLO/SLI definitions, error budgets, reliability engineering principles, and familiarity with secrets management solutions (Azure Key Vault, Hashi Corp Vault).
- Proven ability to coach and mentor engineering teams with strong communication skills across both…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).