Senior Forward Deployed Engineer (DevOps/SRE
Job in
Pleasanton, Alameda County, California, 94566, USA
Listed on 2026-08-07
Listing for:
LeoForce
Full Time
position Listed on 2026-08-07
Job specializations:
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Job Description & How to Apply Below
Job Description
Experience: Senior Level
Salary: $300,000 - $350,000 per year
Job DetailsResponsibilities:
- Implement and optimize an AI-powered Site Reliability Engineering (SRE) platform to meet customer needs across production and pre-production environments.
- Proactively monitor customer deployments to ensure customers maximize value from the platform.
- Identify latent reliability issues such as misconfigurations, deployment regressions, and scaling challenges within customer environments.
- Recommend best practices for implementing AI-powered SRE solutions.
- Plan, design, build, and maintain highly scalable, reliable, and efficient cloud infrastructure.
- Serve as the customer's technical advocate with internal engineering and product teams.
- Conduct post-incident reviews to identify root causes and implement preventative measures.
- Ensure security best practices are integrated into customer deployments.
- Train customer SRE, Operations, and Platform Engineering teams on platform usage and best practices.
- Lead enterprise migrations from legacy alerting, AIOps, and incident management platforms, including correlation rule migration, phased cutovers, and production go-live execution.
- Design, build, and optimize alert normalization and correlation policies using conditions, regular expressions, field extraction, and customized workflows.
- Integrate the platform with customer operational systems, including ITSM, collaboration, observability, source control, and documentation platforms.
- Validate and continuously improve AI investigation quality by tuning enrichment, root cause analysis accuracy, and investigation workflows.
- Build proactive monitoring for customer deployments to identify issues before they impact customers.
- Own customer-facing project communications, including executive status updates, SLA documentation, escalation management, and implementation tracking.
- Develop long-term technical relationships with senior engineering leadership.
- Own customer implementations from technical discovery through solution design, implementation, user acceptance testing, production go-live, stabilization, and ongoing optimization.
- Translate ambiguous customer requirements into clear technical designs, milestones, acceptance criteria, and execution plans.
- Design and implement AI-powered investigation and automation workflows with appropriate guardrails, governance, deterministic fallbacks, and human oversight.
- Develop reusable deployment modules, reference architectures, implementation guides, and operational runbooks to accelerate future deployments.
- Define customer success metrics, establish baselines, measure operational improvements, and demonstrate business value through KPIs such as MTTR reduction and operational efficiency.
- Capture customer feedback and recurring implementation learnings to influence future product development.
- Foster a culture of continuous improvement and technical excellence.
Qualifications:
- Customer-focused with deep empathy for SRE, Dev Ops, Platform Engineering, and IT Operations teams.
- Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
- 6+ years of experience in Site Reliability Engineering, Dev Ops, Platform Engineering, or similar infrastructure-focused roles, including technical leadership or end-to-end customer delivery.
- Experience in Forward Deployed Engineering, Solutions Engineering, Technical Customer Success, or Professional Services is highly preferred.
- Strong programming experience in at least one language such as Python, Go, or Java.
- Hands-on experience with public cloud platforms (AWS, Azure, or Google Cloud Platform).
- Strong knowledge of Kubernetes, Infrastructure as Code (Terraform, Cloud Formation, Ansible), and CI/CD pipelines.
- Practical experience using Generative AI and machine learning technologies to improve engineering productivity.
- Experience with observability platforms, ITSM systems, and incident management tools, including systems integration and data mapping.
- Strong troubleshooting, analytical, and debugging skills, including alert correlation, normalization, and regular expression development.
- Excellent written and verbal…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×