×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer - SDN

Remote / Online - Candidates ideally in
San Francisco, San Francisco County, California, 94199, USA
Listing for: Lambda
Full Time, Remote/Work from Home position
Listed on 2026-08-15
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

Note:

This position requires presence in our San Francisco/San Jose/Bellevue office location 4 days per week;
Lambda’s designated work from home day is currently Tuesday.

Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.

What You'll Do
  • Operate and scale Lambda’s multi-tenant cloud networking platform and SDN infrastructure
  • Operate and improve Kubernetes-based control plane services and dataplane software running on SmartNICs
  • Develop tooling and automation to reduce operational toil and improve reliability
  • Collaborate with software, platform, and networking teams to improve service reliability and deployment workflows
  • Deploy and maintain network monitoring, observability, and management tools
  • Improve deployment safety through CI/CD pipelines, Git Ops workflows, testing, and progressive rollouts
  • Drive operational excellence through observability, incident management, capacity planning, postmortems, and participation in the on-call rotation
You
  • Have 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role
  • Have experience operating and supporting large-scale distributed systems in production
  • Have experience with Kubernetes application lifecycle management, upgrades, troubleshooting, and production operations
  • Have experience participating in on-call rotations and incident response
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary