Sr. SRE – AI Platforms
Listed on 2026-09-01
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Senior Site Reliability Engineer (SRE) - AI Platforms
We’re looking for a Senior SRE, AI Platforms to provide Site Reliability Engineering support for AI platforms and infrastructure environments. This role is focused on platform reliability, production troubleshooting, incident response, observability, and operational automation. The ideal candidate is a senior, hands-on engineer who can troubleshoot complex production environments while driving long-term improvements in platform availability and resilience.
What You’ll Be Doing- Operate and improve the reliability of AI platform services, cluster dependencies, and shared infrastructure.
- Lead and support incident triage involving Kubernetes, Linux, storage, networking, scheduling platforms, job orchestration, and dependency failures.
- Define and improve SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident corrective actions.
- Analyze recurring failures and turn manual operational processes into automation and preventative controls.
- Build observability across systems and services using metrics, logs, traces, and event correlation.
- Troubleshoot performance and availability issues impacting AI/ML training workloads, inference services, and internal platforms.
- Partner with infrastructure and validation teams to improve production readiness and change-management safety.
- Drive operational reviews, resilience testing, and readiness assessments.
- 7+ years of SRE, production operations, or reliability-focused infrastructure engineering experience.
- Strong hands-on expertise with Linux, Kubernetes, networking, and distributed systems.
- Experience building observability, monitoring, alerting, and incident-response workflows.
- Strong scripting and automation experience with Python, Bash, Go, or similar technologies.
- Experience with incident management, root cause analysis, and post-incident remediation.
- Ability to automate repetitive operational processes and reduce production support overhead.
- Strong troubleshooting skills across complex, highly available production environments.
- Excellent communication and ability to work across technical teams.
- AI platforms, machine learning infrastructure, or large-scale HPC environments.
- Prometheus, Grafana, ELK Stack, Open Search, Loki, and Pager Duty.
- Defining and managing SLIs, SLOs, and error budgets.
- Supporting both batch and service-based workloads.
- Cloud and data center infrastructure.
- Logging, tracing, infrastructure monitoring, and service reliability tooling.
Mainz Brady Group is a technology staffing firm with offices in California, Oregon, Washington, and Texas. We specialize in Information Technology and Engineering placements on a Contract, Contract-to-Hire, and Direct Hire basis. Mainz Brady Group is a recipient of multiple annual Excellence Awards from the Tech Serve Alliance, the leading association for IT and engineering staffing firms in the U.S.
Mainz Brady Group is an Equal Opportunity Employer. We are committed to Diversity & Inclusion and incorporate non-discrimination best practices in all our staffing processes.
Mainz Brady Group does not discriminate based on race, color, religion, sex, sexual orientation, gender identity, gender expression, age, disability, or any other protected class.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).