Senior Engineer – ML
Listed on 2026-08-01
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Why choose us?
Are you ready to take the next step in your career? Join us for an exciting opportunity at Albertsons Companies, where innovation and customer service go hand-in-hand! At Albertsons Companies, we are looking for someone who is not just seeking a job, but someone who wants to make an impact. In this role, you'll have the opportunity to lead, innovate, and contribute to the growth of a company that values great service and lasting customer relationships.
This position offers the chance to work in a fast-paced, dynamic environment that is constantly evolving.
Are you ready to take the next step in your career? Join us for an exciting opportunity at Albertsons Companies, where innovation and customer service go hand-in-hand! At Albertsons Companies, we are looking for someone who is not just seeking a job, but someone who wants to make an impact. In this role, you'll have the opportunity to lead, innovate, and contribute to the growth of a company that values great service and lasting customer relationships.
This position offers the chance to work in a fast-paced, dynamic environment that is constantly evolving.
This role is an individual contributor position responsible for designing, developing, fine-tuning, and operationalizing AI/ML capabilities for the AIOps platform within the Observability product. The candidate will work closely with the Lead Engineer, SRE teams, Platform team, Data Ingestion team, Platform Dev Ops team, Visualization team, and other portfolio teams.
As part of the AIOps Platform team, you will design and build intelligent systems that improve observability, incident response, and operational efficiency. This includes developing machine learning models for forecasting, anomaly prediction, alert classification, event intelligence, and causal analysis, as well as building AI agents and multi-agent workflows for RCA summarization, investigative assistance, and SRE productivity use cases.
This position will be based out of Phoenix, Arizona or Pleasanton CA.
Main Responsibilities- Design, develop, and product ionize AI/ML capabilities for the AIOps platform to support intelligent observability and operational decision-making.
- Build machine learning systems for time-series forecasting, anomaly prediction, incident prediction, alert classification, noise reduction, and event correlation.
- Develop causal ML and statistical inference solutions to identify likely root causes, dependency impacts, and relationships across systems and services.
- Create and fine-tune models for incident intelligence use cases such as forecasting service degradation, capacity risk prediction, alert prioritization, and anomaly explanation.
- Design feature pipelines and model training workflows using telemetry, log, metric, trace, topology, and incident data.
- Build intelligent RCA summarization capabilities using LLMs and agentic frameworks such as Lang Chain and Lang Graph.
- Develop AI agents and multi-agent systems for use cases such as Multi-Agent RCA, SRE Assistant, remediation guidance, incident triage, and operational knowledge retrieval.
- Design prompt orchestration, reasoning workflows, retrieval pipelines, tool usage patterns, and memory/context handling for AI agents.
- Integrate AI/ML services with observability platforms, event systems, knowledge bases, CMDB, incident management tools, and automation platforms.
- Collaborate with platform and engineering teams to build scalable model-serving and agent-serving architectures.
- Define and implement evaluation frameworks for model quality, agent effectiveness, hallucination reduction, relevance, and operational usefulness.
- Ensure AI/ML systems are scalable, reliable, explainable, and aligned with enterprise security, governance, and responsible AI practices.
- Build and maintain APIs and microservices for model inference, online scoring, batch predictions, and agent orchestration.
- Partner with SREs, observability engineers, and product stakeholders to translate operational pain points into ML and AI-driven solutions.
- Continuously improve model performance, feature quality, inference latency, agent reliability, and business impact through…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).