More jobs:
DevOps Manager
Job in
Los Gatos, Santa Clara County, California, 95032, USA
Listed on 2026-07-25
Listing for:
GCS Recruitment
Full Time
position Listed on 2026-07-25
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Project Manager, Systems Engineer
Job Description & How to Apply Below
About the Role
We are seeking an experienced Senior Dev Ops Manager to lead a high-performing Technical Operations team responsible for ensuring the reliability, scalability, and performance of a large-scale cloud infrastructure. This leadership role combines technical expertise with people management, driving operational excellence, infrastructure automation, and cross-functional collaboration across engineering, security, and product teams.
The ideal candidate is an experienced Dev Ops or Site Reliability Engineering (SRE) leader with a strong background in cloud infrastructure, incident management, infrastructure automation, and team development.
Key Responsibilities- Lead, mentor, and develop a high-performing Technical Operations/Dev Ops team.
- Ensure 24/7 operational stability, availability, and reliability of production infrastructure.
- Drive operational excellence through continuous improvement initiatives.
- Lead incident response, major incident management, and post-incident reviews (RCA).
- Establish and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
- Oversee infrastructure automation using Infrastructure as Code (IaC) principles.
- Manage Kubernetes-based container platforms and cloud infrastructure.
- Improve deployment processes through CI/CD and automation.
- Collaborate with Product, Engineering, Security, and Infrastructure teams on strategic initiatives.
- Drive disaster recovery, resiliency, and business continuity planning.
- Develop monitoring, alerting, and observability strategies.
- Present operational metrics and project updates to senior leadership.
- Ensure compliance with security standards, audits, and regulatory requirements.
- 8+ years of experience in Dev Ops, Technical Operations, Infrastructure Engineering, or Site Reliability Engineering.
- 4+ years of experience managing technical operations, infrastructure, or SRE teams.
- Strong leadership, mentoring, and people management skills.
- Experience managing production cloud infrastructure in AWS and/or GCP.
- Hands-on experience with Kubernetes, Docker, and container orchestration.
- Strong experience with Terraform and Infrastructure as Code (IaC).
- Experience implementing CI/CD pipelines and deployment automation.
- Strong understanding of Linux systems administration.
- Experience with incident management, on-call operations, and production support.
- Knowledge of networking fundamentals, cloud security, and infrastructure best practices.
- Experience with Agile methodologies, Kanban, Jira, or similar project management tools.
- Excellent communication, stakeholder management, and executive presentation skills.
- Experience implementing Site Reliability Engineering (SRE) practices.
- Experience with monitoring and observability platforms such as Prometheus, Grafana, APM tools, and centralized logging solutions.
- Experience with disaster recovery planning and multi-region infrastructure.
- Knowledge of compliance frameworks including SOC 2 and other security standards.
- Experience supporting enterprise-scale cloud platforms.
- Cloud certifications such as AWS Solutions Architect or Google Cloud Professional certifications.
- Experience with enterprise networking, telecommunications, wireless networking, or IoT environments.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
- Google Cloud Platform (GCP)
- Amazon Web Services (AWS)
- Kubernetes
- Docker
- Terraform
- Linux
- CI/CD Pipelines
- Git
- Prometheus
- Grafana
- Application Performance Monitoring (APM)
- Infrastructure as Code (IaC)
- Jira
- Agile / Kanban
- Networking & Cloud Security
- Opportunity to lead and mentor a high-performing Dev Ops and Technical Operations team.
- Work on large-scale, cloud-native infrastructure supporting enterprise applications and mission-critical services.
- Drive strategic initiatives in infrastructure automation, cloud modernization, and Site Reliability Engineering (SRE).
- Gain hands-on experience with cutting-edge technologies including Kubernetes, Terraform, AWS/GCP, CI/CD, and Infrastructure as Code (IaC).
- Influence technical direction and collaborate with cross-functional engineering, security, and product teams.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×