Senior Site Reliability Engineer Avaloq Fort Lauderdale, Florida,
Listed on 2026-09-24
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support, Cybersecurity
This job is with Avaloq, an inclusive employer and a member of my Gwork – the largest global platform for the LGBTQ+ business community. Please do not contact the recruiter directly.
Company Description
Founded and headquartered in Switzerland, Avaloq is continuously expanding its global footprint with around 2,500 colleagues in 10 countries, and more than 170 clients in 35 countries. We are an industry-leading provider of wealth management technology and services for financial institutions around the world, including private banks and wealth managers, investment managers, as well as retail and neo banks. Our research led approach and continual innovation is powered by the passion and creativity of our colleagues.
We are always looking for talented people to join us on our mission to orchestrate the financial ecosystem and democratize access to wealth management. Avaloq offers the opportunity to work closely with some of the world’s leading financial institutions as we jointly develop and shape careers. Championing a collaborative, supportive and flexible work environment empowers our colleagues to reach their full potential.
Job Description
Avaloq's R&D Lab is building a SaaS, API-first, composable banking platform. As a Senior Site Reliability Engineer you will help build the reliability practice behind it.
This is amongst the first technical roles we are hiring for our Fort Lauderdale site. You will partner with our Senior SRE in Zurich and work alongside our platform engineers and product teams.
The first three to six months:Some of what you can expect early on:
- Learning our AWS environment, our serverless platform, and how our product teams build and release
- Contributing to the direction on observability tooling, together with the platform engineers, the SRE team, and our developers
- Working with the Zurich team on the incident and on-call operating model
- Taking on increasing responsibility as you build context, always in alignment with the Zurich team
- Define and evolve SLIs, SLOs, and error budgets with product teams, and use them to drive reliability decisions
- Collaborate on the design of our observability approach for a distributed serverless system, covering metrics, logs, and traces
- Build the incident response practice with us: on-call model, escalation, blameless post-mortems, and the loop that turns findings into hardening work
- Build reliability automation that detects and remediates issues before they reach clients
- Improve CI/CD pipelines and deployment automation on Git Hub Actions to reduce operational toil and release risk
- Work with product teams on resilient design, capacity planning, and progressive delivery approaches such as canary and blue-green
- Contribute to disaster recovery design and testing for client-facing environments
- Partner with Security and Compliance so that operational practice holds up to regulatory and audit expectations
- Help us define and implement operational readiness for client go-live
- Mentor colleagues and support a culture of shared operational ownership
Our stack:
Serverless-first on AWS:
Lambda, DynamoDB, SQS, S3, and Bedrock. Terraform and Open Tofu for infrastructure, Git Hub Actions for CI/CD. Product teams work in Rust, Type Script, Python, Vue, and Angular. We do not run containers.
Qualifications
- 5+ years in Site Reliability Engineering, Dev Ops, or production operations for distributed cloud systems, with substantial hands-on AWS experience
- Practical experience defining and working with SLIs, SLOs, and error budgets
- Strong incident response background, including on-call, triage under pressure, and post-mortem practice
- A clear point of view on observability for distributed systems, and the reasoning behind it
- Solid automation and scripting…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).