SRE Engineer with ML Ops and Arize product
Listed on 2026-08-17
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
AWS Cloud Platform Expertise
EC2, EKS, ECS, Lambda, Cloud Watch, SNS, SQS, Event Bridge for highly available and scalable services.
Arize AI Observability & MonitoringModel performance monitoring, drift detection, evaluation analytics, and AI/LLM observability.
Site Reliability Engineering (SRE)SLI/SLO definition, error budgets, reliability improvement, service resilience, and uptime management.
Incident Management & RCAMajor incident management (MIM), outage tracking, root cause analysis, problem management, and MTTR reduction.
Failure Analysis & Risk AssessmentFMEA, risk quantification, reliability assessments, and proactive mitigation of platform failures.
Testing & Validation EngineeringScenario testing, regression testing, impact analysis, release validation, and upstream change testing.
Monitoring, Alerting & AutomationCloud Watch, Grafana, Prometheus, Pager Duty, automated notifications, dashboards, and operational metrics.
Dev Ops & MLOps PracticesKubernetes, Terraform, CI/CD pipelines, Python scripting, AI/ML platform operations, and LLM reliability optimization.
Preferred TechnologiesAWS, Arize AI, Kubernetes (EKS), Cloud formation, Python, Cloud Watch, Grafana, Prometheus, Pager Duty, Git Hub Actions/Jenkins.
Diverse Lynx LLC is an Equal Employment Opportunity employer. All qualified applicants will receive due consideration for employment without any discrimination. All applicants will be evaluated solely on the basis of their ability, competence, and their proven capability to perform the functions outlined in the corresponding role. We promote and support a diverse workforce across all levels in the company.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).