Site Reliability Engineer
Irvine, Orange County, California, 92713, USA
Listed on 2026-10-08
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Title :
Site Reliability Engineer
Duration:
Initial 3–6 months, C2H
Location:
Irvine, CA
- Onsite, 5 days per week
- Updated resume
- Candidate location & onsite confirmation
- Years of experience in AWS, Kubernetes, Terraform, CI/CD
- SRE/Dev Ops experience summary
The SRE will embed directly with product teams to support platform reliability, troubleshoot issues, and act as the bridge between development, platform engineering, and cloud infrastructure. This is not a coding-heavy role and focuses on systems understanding, troubleshooting, Kubernetes operations, cloud infrastructure, CI/CD pipelines, and production stability. The team is newly formed and supports product-facing platform needs after splitting from core cloud/Dev Ops engineering.
SeniorityLevel / Target Experience
Mid-level (4–6 years). Not junior; not senior/staff-level.
Top 3–5 Technical SkillsBonus (Not Required):
- Kafka / Streaming
- MLOps / AI infrastructure
- Event-driven or ML-driven platform experience
- Platform ownership post-vendor POC
Immediate need; start as soon as they identify a strong candidate.
Interview Process3 Rounds
- Bachelor’s degree preferred but equivalent experience accepted
- Strong troubleshooting skills across distributed systems
- Must be comfortable in a high-change, incident-driven environment
- Will participate in on-call support and incident response
- Role sits on a new Product Platform team (separate from core Dev Ops/cloud)
- Focus is platform reliability, not feature development
- Must be highly organized, detail-oriented, and able to work independently
Role Overview
As a Site Reliability Engineer, you will be immersed in a high-performing and frequently challenged team responsible for the reliability, scalability, and operational excellence of PDS’s product platform. You will work and consult closely with engineering and product teams to design, build, and operate resilient platform services, implement automation and reliability frameworks, and ensure the availability and performance of shared platform capabilities.
You will also be responsible for the ongoing operation, monitoring, and continuous improvement of platform infrastructure and services that support all product teams.
- Cloud Architecture & Infrastructure Engineering
- Architect and maintain multi-cloud infrastructure (AWS and GCP and Azure) to support enterprise-scale healthcare operations.
- Implement Infrastructure-as-Code (IaC) using Terraform, Helm charts, and Cloud Formation to automate resource provisioning and ensure consistency across environments.
- Configure, manage, and deploy Kubernetes clusters and cloud native tooling according to best practices for scaling, resiliency, and reliability.
- Manage and optimize multi-cluster Kubernetes environments, utilizing Istio service mesh for advanced traffic management, service discovery, and observability.
- Enforcing security standards and SAST, DAST, code quality scans on the Kubernetes cluster and application pipelines and containers and remediating any security findings related to the platform.
- Design solutions that are aligned with business value in regard to TCO and ROI.
- Design cross‑region disaster recovery strategies and rolling deployment…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).