Senior AI Observability engineer
Listed on 2026-09-12
-
IT/Tech
AI Engineer (Applied/Software)
In your career, let’s prove what’s possible.
At Lam Research, we create equipment that drives technological advancements in the semiconductor industry. Our innovative solutions enable chipmakers to power progress in nearly all aspects of modern life, and it takes each member of our team to make it possible.
Across our organization, our employees come to work and change the world. We take on the toughest challenges with precision and accuracy. We push for the next big semiconductor breakthrough. We lead the way in one of the most critical and fast-moving industries on the planet. And we do it together, with deep connections and limitless collaboration.
The impact we have on the world is made possible by focusing on our people. So we recognize and celebrate our teams' achievements. We strive to create an inclusive and diverse culture where everyone's contribution and voice has value. We evaluate and evolve our offerings, so our people receive the support and empowerment to do meaningful things for their lives, careers, and communities.
Because at Lam, we believe that when people are the priority and they’re inspired to unleash the power of innovation for a better world together, anything is possible.
Senior AI Observability engineerDate:
Aug 19, 2026
Fremont, CA, US, 94538
Worker Category:
On‑site Flex
We are seeking a hands‑on Senior AIOps Reliability Engineer to build the AI‑native operations layer for our hybrid enterprise estate. You will design and ship LLM‑based agents, retrieval‑augmented knowledge pipelines, and machine‑learning anomaly detection that find, triage, and remediate incidents across public cloud, private cloud, and on‑premises data centers, all resting on strong SRE and network engineering fundamentals.
Our environment spans Azure, AWS, and Google Cloud alongside on‑premises data centers, colocation sites, and manufacturing and HPC facilities, so cloud‑agnostic design and hybrid network fluency matter more than depth in any single provider. This is a deep individual‑contributor role: you will write the code, instrument the telemetry, tune the models, and own the reliability of both the infrastructure and the AI systemsoperatingon it.
The impact you’ll makeJoin Lam as an IT Engineer, where you'll be at the forefront of designing, analyzing, and implementing applications and systems that form the foundation of our infrastructure. As a crucial member of our IT team, you'll contribute your technical assistance and guidance to projects for various systems and infrastructures. Acting as a technical liaison, you'll address complex business problems with automated systems solutions.
Your expertise will be instrumental in driving Lam's commitment to innovation and efficiency.
AI and Agentic Operations
- Build agentic AI workflows using LLM agents, tool and function calling, and orchestration frameworks such as Lang Graph, Semantic Kernel, Auto Gen, or the Model Context Protocol, applied to autonomous fault detection, triage, and remediation.
- Develop the AIOps intelligence layer: time‑series anomaly detection, dynamic baselining, alert deduplication and correlation, event clustering, and predictive failure and capacity forecasting across infrastructure, network, and application telemetry from both cloud and on‑premises sources.
- Engineer the retrieval knowledge fabric by chunking, embedding, and indexing runbooks, post‑mortems, architecture documents, CMDB and Service Now records into a vector store, then tuning retrieval quality against measurable evaluations.
- Ship AI‑assisted incident response: automated summarization, root‑cause hypothesis generation, blast‑radius analysis, and telemetry‑grounded draft post‑mortems wired into the paging…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).