Site Reliability Engineer (DevOps Engineer
Job in
Saratoga, Santa Clara County, California, 95071, USA
Listed on 2026-08-08
Listing for:
E-Space
Full Time
position Listed on 2026-08-08
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, AWS
Job Description & How to Apply Below
- Design, deploy, and maintain highly-scalable, highly-available software systems in AWS
- Architect and manage containerized applications on Amazon EKS with focus on reliability and performance
- Build and maintain Infrastructure as Code using Terraform for AWS cloud resources
- Develop and optimize CI/CD pipelines for automated testing, deployment, and rollback capabilities
- Implement comprehensive monitoring, alerting, and observability solutions using Cloud Watch, Prometheus, and Grafana
- Ensure system reliability through SLI/SLO definition, error budgets, and incident response procedures
- Collaborate directly with engineering teams to optimize application deployment and operations
- Manage deployments and scaling strategies to support mission-critical operations
- Automate and enforce cloud security, governance, and compliance controls
- Participate in on-call rotation and lead incident response for production level systems
- Experience with database operations and scaling (RDS, Aurora, or similar)
- Proficiency in Python and Bash scripting for automation
- 5+ years of experience in SRE, Dev Ops, or Platform Engineering roles
- Deep experience with Kubernetes, EKS, Helm, and container orchestration
- Strong CI/CD pipeline development and management experience (Bitbucket preferred)
- Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack)
- Proven experience designing and operating mission-critical, highly-available systems within AWS
- Knowledge of capacity planning and performance optimization
- Advanced proficiency in Infrastructure as Code using Terraform (Open Tofu)
- AWS Solutions Architect Professional, Certified Kubernetes Administrator (CKA), or equivalent expertise
- Experience with incident management and post-mortem processes
- Experience with Git Ops workflows and tools (ArgoCD, Flux)
- Experience with chaos engineering and disaster recovery planning
- Knowledge of service mesh technologies (Istio, Linkerd)
- Experience with Zero Trust Networking (ZTNA) or VPN solutions
- Background in aerospace, defense, or other mission-critical industries
- Strong intellectual curiosity and commitment to continuous learning
- Exceptional attention to detail and an ownership mentality
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×