Senior Site Reliability Engineer
Job Description & How to Apply Below
We’re looking for an incredible Senior Site Reliability Engineer to join our SRE team. We aim to make reliability, security, and speed reinforce one another so that the platform becomes the engine of Relay’s growth. Your love of making high-impact decisions daily and desire to help shape the future of Relay is going to be crucial.
What You’ll Be Doing
Join the team building and owning our production infrastructure and CICD pipelines (AWS, Kubernetes, PostgreSQL databases, Terraform, Terragrunt, Github Actions)
Review infrastructure change requests, and triage & fix high-risk security and privacy issues in infrastructure components
Build monitoring systems to dynamically assess the infrastructure health
Improve our data repositories (DB, data warehouse, datalake) posture: engine upgrades, zero-downtime migrations, privacy taggings
Partner with product teams to balance feature delivery with reliability constraints
Establish and maintain error budgets for core services
Provide guidance and mentoring for the rest of the team and help evolve Relay into a world-class security-oriented organization
Write runbooks and coordinate gamedays
Shape and mature Relay’s service ownership model by working closely with engineering teams to clarify operational responsibilities
Collaborate in leading incident response by driving fast mitigation, clear communication, and structured decision-making
Who You Are
You have experience as a site reliability engineer working with these technologies: AWS, Kubernetes, Datadog, Git Hub, and Git Hub Actions
You have experience owning observability initiatives (logging pipelines, Software Catalog, monitoring strategies, and incident management tooling)
You have various levels of experience with Terraform, Terragrunt, Node.js, Typescript
You have hands-on experience managing and optimizing databases such as Aurora RDS, PostgreSQL, DynamoDB, and Elasti Cache
You have a strong security and operations mindset. We are looking for someone to help us continue building security into every aspect of our work - and is ready to be on-call for production issues
You are a team player. Our team is lean, high-impact, and deeply collaborative - we want someone who is always willing to pitch in and isn’t afraid to ask for help
You are curious. You keep yourself on the bleeding edge of infrastructure best practices
Bonus Points
Show us your home lab
Show us your Github profile!
Fintech/regulatory experience;
Experience working in compliant environments such as SOC2 or PCI
Experience driving large reliability initiatives across your company
You’ve joined a company at its early stages and have seen it through scale
You have experience working in a fintech startup
The Interview Process
Stage 1: A 45-minute Google Meet video call with a member of our Talent team
Stage 2: A 45-minute Google Meet video call with the Hiring Manager
Stage 3: A 45-minute Google Meet video call with a member of our leadership team
Stage 4: A take-home case study followed by a 60-minute in-person presentation to members of our SRE team.
Our Compensation Approach
We believe Relayers should feel rewarded for the impact they have on our mission and growth. Compensation follows impact. As impact increases, compensation grows, and we do not limit compensation changes to a once-a-year review cycle.
The annual salary…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×