×
Register Here to Apply for Jobs or Post Jobs. X

Senior Manager, Reliability Engineering & AIOps

Job in Fremont, Alameda County, California, 94537, USA
Listing for: Lam Research Salzburg GmbH
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 137000 - 287000 USD Yearly USD 137000.00 287000.00 YEAR
Job Description & How to Apply Below

In your career, let’s prove what’s possible.

At Lam Research, we create equipment that drives technological advancements in the semiconductor industry. Our innovative solutions enable chipmakers to power progress in nearly all aspects of modern life, and it takes each member of our team to make it possible.
Across our organization, our employees come to work and change the world. We take on the toughest challenges with precision and accuracy. We push for the next big semiconductor breakthrough. We lead the way in one of the most critical and fast-moving industries on the planet. And we do it together, with deep connections and limitless collaboration.
The impact we have on the world is made possible by focusing on our people. So we recognize and celebrate our teams’ achievements. We strive to create an inclusive and diverse culture where everyone’s contribution and voice has value. We evaluate and evolve our offerings, so our people receive the support and empowerment to do meaningful things for their lives, careers, and communities.
Because at Lam, we believe that when people are the priority and they’re inspired to unleash the power of innovation for a better world together, anything is possible.

Senior Manager, Reliability Engineering & AIOps

Date:
Aug 27, 2026

Location:

Fremont, CA, US, 94538

Worker Category:
On-site Flex

The group you’ll be a part of

You will join the Reliability Engineering team within Infrastructure Platform Engineering. The group keeps Lam's global infrastructure estate available and recoverable across Azure, AWS, GCP, compute, storage, network, and high-performance computing, supporting engineering and operations teams in the US, Japan, Singapore, Malaysia, India, and Korea.

The impact you’ll make

As Senior Manager of Reliability Engineering& AIOps, you lead the team that keeps critical infrastructure running and proves it is ready for the next failure. In this role, you will directly contribute to the availability of the systems Lam's engineering, manufacturing, and business teams depend on every day, and you will build the automation that makes outages rare, short, and unremarkable.

What you’ll do
  • Lead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on.
  • Set the reliability strategy: define the service level objective program, publish an error-budget policy, and drive adoption across platform and service teams.
  • Build and run a follow-the-sun on-call and response model across six regions, with clean handoffs and one consistent set of runbooks and severity definitions worldwide.
  • Own the incident management and paging platform end to end, including services, schedules, escalation policies, and routing, configured as code and tuned so alerts fire on real risk rather than noise.
  • Serve as incident commander on major incidents, own executive and stakeholder communications, and lead blameless postmortems with tracked follow-up.
  • Own disaster recovery strategy and execution across Azure, AWS, GCP, and core infrastructure platforms, including service-tier recovery objectives, backup and restore validation, failover readiness, DR certification, runbook governance, and recurring exercises measured against RTO and RPO targets.
  • Lead capacity planning and performance engineering across Azure, AWS, GCP, compute, storage, network, and HPC platforms, using demand forecasting, utilization trends, growth modeling, and automation to prevent capacity risk and reduce manual operational work.
  • Define and drive AI Ops requirements for reliability engineering across Azure, AWS, and GCP, including Microsoft Copilot, Cursor, Git Hub Copilot, and LLM-based operational workflows for incident triage, runbook generation, knowledge retrieval, root-cause…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary