More jobs:
Distinguished Software Engineer – AI/ML Engineer
Job in
Sunnyvale, Santa Clara County, California, 94087, USA
Listed on 2026-09-09
Listing for:
Jobtailor
Full Time
position Listed on 2026-09-09
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Cloud Engineer - Software, DevOps
Job Description & How to Apply Below
- Lead technical development of next-generation agentic AI systems and intelligent automation solutions for mission-critical reliability, scalability, and operational excellence
- Architect and implement machine learning platforms and autonomous agents for change management, monitoring, prediction, and automated issue resolution
- Design and implement multi-agent orchestration platforms for change management, capacity planning, and performance optimization
- Build intelligent observability and monitoring platforms using ML-driven anomaly detection, predictive analytics, and autonomous resolution
- Develop self-healing infrastructure platforms that predict, prevent, and automatically remediate system issues
- Design and build tools improving latency, availability, scalability, and change management
- Engineer reliability using metrics and measurements across domains
- Enable system scaling through technical solutions, automation, and process optimization
- Build failure-prevention tools and automation for mission-critical services
- Enhance instrumentation for cohesive end-to-end system health visibility
- Architect and implement fault-tolerant systems across hybrid cloud infrastructure
- Reduce MTTD and MTTR through intelligent automation and predictive capabilities
- Partner with service owners to define SLA breach detection and change-related anomaly handling
- Troubleshoot and analyze large-scale distributed systems
- Deliver autonomous reliability solutions using machine learning, NLP, and computer vision
- Drive development of MLOps and AIOps platforms for continuous learning, deployment, monitoring, and autonomous optimization
- Implement CI/CD pipelines with automated validation, deployment, rollback, and observability
- Build reusable reliability infrastructure, intelligent monitoring platforms, and developer productivity tools
- Provide technical mentorship and thought leadership through code reviews, design discussions, and knowledge sharing
- Bachelor’s degree in computer science, computer engineering, computer information systems, software engineering, or related area and 6 years’ experience in software engineering or related area; OR 8 years’ experience in software engineering or related area
- 12+ years of hands-on experience in Reliability Engineering, AI/ML Engineering, or Platform Engineering
- Proven record as a senior individual contributor influencing architecture and driving technical excellence across large organizations
- Deep experience operating mission-critical systems
- Expertise in MTTD, MTTR, availability, change management, model performance, and autonomous system reliability
- Expert-level AI/ML engineering experience, including Tensor Flow, PyTorch, and large-scale production ML deployments
- Advanced experience with agentic AI systems, multi-agent frameworks, autonomous decision-making systems, LLM-based agents, and agent orchestration platforms
- Comprehensive Reliability Engineering expertise, including Incident, Problem, and Change Management and performance and capacity engineering for AI/ML systems
- Expert-level cloud engineering experience with Azure, GCP, or AWS
- Experience with Kubernetes, Docker, serverless architectures, and cloud-native AI services
- Deep observability experience across distributed tracing, metrics, logs, APM, and AI-driven anomaly detection
- Strong platform engineering background including infrastructure as code, service mesh architectures, API gateways, and self-service developer platforms
- Preferred:
Master’s degree and 4 years' experience in software engineering or related area - Preferred: MLOps and model lifecycle management using MLflow, Kubeflow, or Seldon
- Preferred: NLP and computer vision expertise
- Preferred:
Edge computing and distributed systems experience - Preferred:
Kafka or Pulsar - Preferred:
Chaos engineering, fault injection, and performance optimization - Preferred:
Open-source contributions in reliability, observability, or infrastructure automation - Knowledge of WCAG 2.2 AA standards, assistive technologies, and digital accessibility best practices is preferred
Demonstrates expertise in Reliability Engineering, AI/ML Engineering, and Platform Engineering, with a focus on architecting and implementing autonomous systems, intelligent automation solutions, and machine learning platforms. Proficient in cloud engineering and observability practices, ensuring mission-critical system reliability and performance optimization.
Highest-signal resume keywords- Reliability Engineering
- AI/ML Engineering
- Cloud Engineering
- MLOps
- Multi-Agent Frameworks
- Machine Learning
- Tensor Flow
- Py Torch
- Kubernetes
- Docker
- CI/CD Pipelines
- Infrastructure as Code
- Predictive Analytics
- Anomaly Detection
- Change Management
- Technical Mentorship
- Thought Leadership
- Incident Management
- Problem Management
- Change Management
- Performance Engineering
- Capacity Engineering
- Azure
- GCP
- AWS
- MLflow
- Kubeflow
- Seldon
- Kafka
- Pulsar
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×