Manager, Engineering - Dev Ops/SRE; Hybrid
Listed on 2026-07-20
-
Software Development
About the Role
At Crowd Strike, Site Reliability Engineering (SRE) is at the forefront of ensuring the reliability and scalability of our cloud-native security platform. As an SRE Engineering Manager, you will lead a team responsible for managing the complex challenges of scale unique to Crowd Strike, leveraging your expertise in software engineering, systems design, and automation. You will play a critical role in ensuring that our services maintain the highest levels of reliability, uptime, and performance, meeting the needs of our customers while continuously improving our systems.
You will have the opportunity to work on meaningful projects while providing support and mentorship to your team, enabling them to learn, grow, and make a lasting impact in the cybersecurity landscape. The ideal candidate will have hands‑on experience in cloud solutions development, strong leadership skills, and a collaborative approach to working with cross‑functional teams. Given our remote‑first culture, exceptional verbal and written communication skills are essential for effective collaboration with engineering teams and colleagues worldwide.
Prior experience in the security industry is not required for this role.
As an SRE Engineering Manager, you will lead a team responsible for managing the complex challenges of scale unique to Crowd Strike, leveraging your expertise in software engineering, systems design, and automation. You will play a critical role in ensuring that our services maintain the highest levels of reliability, uptime, and performance, meeting the needs of our customers while continuously improving our systems.
You will work on meaningful projects while providing support and mentorship to your team, enabling them to learn, grow, and make a lasting impact in the cybersecurity landscape.
- Proven track record of building, growing, and retaining high‑performing SRE/Platform engineering teams in a fast‑paced, high‑growth environment.
- 10+ years of software engineering experience with significant focus on reliability engineering, platform infrastructure, and production operations at scale.
- 3+ years of hands‑on management experience overseeing SRE/Platform engineering teams, including incident command and reliability ownership.
- Deep understanding of SRE principles including SLOs, SLAs, SLIs, error budgeting strategies applied to large‑scale distributed systems.
- Driving system reliability by blending software engineering principles with AI‑driven automation, moving from reactive firefighting to proactive, automated operations.
- Proficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi‑region architecture, cloud‑native reliability patterns, and security‑first cloud design.
- Proven experience owning reliability for high‑throughput distributed systems processing millions of events per second, including capacity planning, traffic management, and load shedding strategies.
- Strong incident management background including leading major incident response, facilitating blameless post‑mortems, and driving systemic reliability improvements.
- Demonstrated ability to build, operationalize, and maintain highly scalable, security‑critical systems with zero tolerance for data loss or downtime.
- Bachelor’s degree in Computer Science or related field, or equivalent work experience.
- Ability to work 2+ days per week in our Sunnyvale Offices.
- Experience operating security platforms, telemetry pipelines, or sensor fleet infrastructure at massive scale (millions of connected endpoints, petabyte‑scale data processing).
- Proficiency in Golang, Crowd Strike’s primary backend language for platform services.
- Experience with Kubernetes at scale managing large cluster fleets, including service mesh technologies (Istio/Linkerd) and container security practices.
- Familiarity with high‑throughput data streaming platforms such as Apache Kafka and Apache Flink for real‑time event processing.
- Experience with hybrid cloud environments spanning cloud and on‑premise data center infrastructure including multi‑cloud failover strategies.
- Advanced observability experience including Prometheus,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).