More jobs:
Site Reliability Engineer (SRE
Job in
Sunnyvale, Santa Clara County, California, 94086, USA
Listed on 2026-08-05
Listing for:
Ace Stack
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Network Engineer
Job Description & How to Apply Below
Site Reliability Engineer (SRE)
Location:
Sunnyvale, CA (3x/ week onsite)
Contract
Responsibilities:
- Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure solutions.
- Operate, monitor, and triage all aspects of our production and non-production environments.
- Collaborate on code, infrastructure, design reviews, and process enhancements Evaluate and integrate new technologies to improve system reliability, security, and performance.
- Develop and implement automation to provision, configure, deploy, and monitor services.
- Participate in an oncall rotation providing hands-on technical expertise during service impacting events.
- Contribute to capacity planning, scale testing, and disaster recovery exercises Approach operational problems with a software engineering mindset.
Min
Qualification:
- 5+ years in Infrastructure Ops, Site Reliability Engineering, or Dev Ops focused role.
- BS degree in computer science or equivalent field with 5+ years of experience.
- Knowledge of Linux operating system principles, networking fundamentals, and systems management.
- Demonstrable fluency in at least one of the following languages:
Java, Python, or Go. - Experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment.
- Familiarity with micro-services architecture and container orchestration with Kubernetes.
- Awareness of key security principles including encryption, keys (types and exchange protocols).
- Understanding of SRE principals including monitoring, alerting, error budgets, fault analysis, and automation.
- Strong sense of ownership, with a desire to communicate and collaborate with other engineers and teams.
- Ability to identify and communicate technical and architectural problems, while working with partners and their team to iteratively find solutions.
Role Description s:
We are seeking a Dev Ops Site Reliability Engineer (SRE) with strong experience in containerization orchestration and automation. The ideal candidate will have hands-on expertise in Kubernetes Docker and Python and will be responsible for building scalable infrastructure automating operations and ensuring high availability of production systems.
Key Responsibilities:
- Design deploy and maintain containerized applications using Docker and Kubernetes.
- Build and maintain automated infrastructure and deployment pipelines.
Develop automation scripts and tools using Python.
Manage and optimize Kubernetes clusters in production environments. - Implement CICD pipelines to streamline build| test| and deployment processes.
- Monitor system performance and reliability using observability and monitoring tools.
- Troubleshoot production issues and participate in incident response and root cause analysis.
- Work closely with development teams to improve system reliability and deployment efficiency.
- Implement security and best practices for container and cloud infrastructure.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×