Senior Site Reliability Engineer - SDN
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-08-15
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
Note:
This position requires presence in our San Francisco/San Jose/Bellevue office location 4 days per week;
Lambda’s designated work from home day is currently Tuesday.
Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.
What You'll Do- Operate and scale Lambda’s multi-tenant cloud networking platform and SDN infrastructure
- Operate and improve Kubernetes-based control plane services and dataplane software running on SmartNICs
- Develop tooling and automation to reduce operational toil and improve reliability
- Collaborate with software, platform, and networking teams to improve service reliability and deployment workflows
- Deploy and maintain network monitoring, observability, and management tools
- Improve deployment safety through CI/CD pipelines, Git Ops workflows, testing, and progressive rollouts
- Drive operational excellence through observability, incident management, capacity planning, postmortems, and participation in the on-call rotation
- Have 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role
- Have experience operating and supporting large-scale distributed systems in production
- Have experience with Kubernetes application lifecycle management, upgrades, troubleshooting, and production operations
- Have experience participating in on-call rotations and incident response
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).