×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer, Data Platform​/Multi-Cloud Infrastructure

Job in Austin, Travis County, Texas, 78701, USA
Listing for: Apple
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Position: Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.

Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products  sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands-on support to internal teams, and partnering with developers to make cutting-edge services like Spark, Flink, Airflow, Ray, Notebooks, and LLM-based agent platforms reliable at scale across AWS, GCP, and on-premise Kubernetes.

Responsibilities
  • Operate, monitor, and triage production and non-production environments across the ADP portfolio — data processing, ML/AI, and multi-cloud infrastructure.
  • Participate in a rotating on-call schedule across supported services, including occasional weekday and weekend coverage.
  • Own the operational health of multi-cloud infrastructure as SME — driving reliability, support, and customer guidance for AWS services, EKS clusters, and cross-cloud networking.
  • Provide Slack-based support to internal customers; screen, triage, and resolve service related issues.
  • Debug production incidents involving IAM permission errors, storage quota limits, control-plane/data-plane namespace separation, and cluster-wide disruptions.
  • Partner with dev teams across time zones to onboard new services — understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
  • Maintain and evolve Infrastructure-as-Code (Crossplane, Terraform) and Git Ops (Flux) workflows, troubleshooting state drift and reconciliation issues.
  • Build automation and self-healing tooling that reduces manual toil and scales the team's operational capacity.
  • Identify, escalate, and resolve production issues to protect platform reliability and customer experience.
  • Collaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals.
Minimum Qualifications
  • * Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, Dev Ops, or Infrastructure-focused role.
  • Proficient in Python; working knowledge of Golang a plus.
  • Strong hands-on AWS experience: IAM (roles, policies, permission boundaries, KMS), EKS, RDS, S3, VPC networking/endpoints, autoscaling groups, EBS.
  • Kubernetes administration experience — RBAC, node/pod scheduling, autoscalers, Priority Classes/PDBs, and troubleshooting cluster-wide disruptions.
  • Strong communication skills and composure under pressure during incidents.
  • Solid grounding in SRE principles, with prior on-call or production-support experience.
Preferred Qualifications
  • Experience with Infrastructure-as-Code (Crossplane and/or Terraform), including debugging state drift and composition/controller issues.
  • Experience with Git Ops workflows (Flux or similar) — Helm Repository/reconciliation troubleshooting and Helm chart deployment.
  • Multi-cloud exposure (GCP) — parity and migration scenarios are emerging areas of focus.
  • Experience with Splunk for log pipeline debugging (e.g., fluent-bit).
  • Familiarity with Spark/Flink running on Kubernetes (executor scheduling, node affinity).
  • Comfort with Git Hub PR review workflows in an infrastructure-as-code / Git Ops context.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning — for yourself, your team, and the org.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant At Apple, we believe accessibility is a fundamental human right.

You'll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple's workplace Learn about reasonable accommodations for job applicants Apple accepts applications to this posting on an ongoing basis.

Submit Resume Back to search results See all roles in Austin

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary