Senior Site Reliability Engineer
Listed on 2026-08-31
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, AWS
Senior Site Reliability & Cloud Systems Engineer
This is a hybrid role - 2 days remote and 3 days in the Malvern, PA office.
We are seeking a highly experienced Senior Site Reliability & Cloud Systems Engineer to architect, build, automate, and operate scalable, secure, resilient, and highly available cloud platforms in AWS. This role combines hands-on reliability engineering with cloud architecture and automation expertise, with a strong emphasis on building immutable infrastructure and improving system resilience.
You will play a critical role in evolving our AWS ecosystem into a fully automated, self-service, "push-button" platform, minimizing manual operational intervention while improving reliability, security, performance, scalability, and engineering velocity. You will establish and champion engineering standards for immutable infrastructure, automation, observability, resilience, and operational excellence.
This role is well suited for a senior engineer who thrives at the intersection of SRE, cloud architecture, platform engineering, Dev Ops, and systems engineering, and who can independently drive complex technical initiatives from architecture and design through implementation and production operation.
At Cube Smart, we're intentional about culture. You can experience it everywhere from our mission statement of "genuine care" to our "It's What's Inside That Counts" tagline to calling each other "teammates" rather than employees. This spirit fosters a fun and collaborative environment that has resulted in our rapid growth and being recognized amongst the top in our industry.
Cube Smart's award-winning team is made up of people who genuinely care. Teammates care about our customers and the life events and/or business needs they are facing. Teammates are passionate, responsible and understanding. The Cube Smart team is made up of people who have a can-do attitude, are committed to their own success and the success of the company, and lead by example.
If this sounds like a team and culture that matches your personal values and motivations, we want to hear from you.
ResponsibilitiesReliability, Performance & Production Operations
- Own the reliability, availability, scalability, performance, and operational health of AWS-hosted, Linux-based production platforms and associated lower environments.
- Define and drive SRE practices, reliability standards, SLOs, SLIs, error budgets, and operational maturity across critical systems.
- Architect and implement comprehensive observability using platforms such as Datadog, Amazon Cloud Watch, Prometheus/Grafana, and Pager Duty.
- Establish proactive monitoring, alerting, capacity planning, performance engineering, and predictive reliability practices.
- Lead complex production incident response, including technical triage, mitigation, service restoration, root cause analysis, and executive/stakeholder communication.
- Drive blameless post-incident reviews and ensure corrective actions are translated into measurable reliability improvements.
- Participate in and provide leadership during 24/7 on-call operations, including escalation management for high-severity production incidents.
- Identify systemic reliability risks and proactively eliminate single points of failure, operational bottlenecks, and sources of technical debt.
- Develop and implement disaster recovery, business continuity, backup, restoration, and resilience strategies.
- Perform capacity and performance analysis for distributed applications and infrastructure at scale.
Cloud Architecture & Automation
- Architect and implement highly automated, ephemeral, immutable, and reproducible AWS environments across production and non-production workloads.
- Lead the design of scalable, fault-tolerant, secure distributed systems using AWS Well-Architected principles.
- Establish infrastructure patterns that enable engineering teams to provision and manage environments through self-service and "push-button" automation.
- Eliminate manual infrastructure operations through Infrastructure as Code using Terraform, Ansible, Packer, and related technologies.
- Design and maintain reusable infrastructure modules, automation frameworks, and engineering standards.
- Build and evolve CI/CD and Git Ops workflows using technologies such as Jenkins, Git Hub Actions, Git Lab CI, ArgoCD, and Flux.
- Develop sophisticated automation and operational tooling using Python and Bash.
- Identify opportunities to reduce operational toil through automation, platform engineering, and intelligent operational workflows.
- Establish engineering patterns for immutable infrastructure, automated provisioning, blue/green deployments, canary releases, and automated rollback.
Infrastructure & Platform Engineering
- Architect, deploy, and operate AWS services including:
- EKS, ECS, Fargate, Lambda
- RDS and Aurora PostgreSQL
- Open Search
- Redis and Elasti Cache
- Load balancing, networking, storage, compute, and supporting AWS services
- Design and manage enterprise AWS networking architectures, including Transit Gateways,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).