×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer, AiDP Production Engineering

Job in Austin, Travis County, Texas, 78716, USA
Listing for: Apple Inc.
Full Time position
Listed on 2026-07-18
Job specializations:
  • Software Development
    Cloud Engineer - Software
Salary/Wage Range or Industry Benchmark: 140000 - 170000 USD Yearly USD 140000.00 170000.00 YEAR
Job Description & How to Apply Below

Site Reliability Engineer, AiDP Production Engineering

Austin Metro Area, Texas, United States Software and Services

The Production Engineering team within the AI and Data Platform (AiDP) organization manages a wide array of real-time, near real-time, and batch analytical solutions. These platforms are integral to core business functions across Apple. These include sales, operations, finance, Apple Care, marketing, and services, and are instrumental in driving critical, data-driven decisions. To build these solutions, we leverage a combination of proprietary and leading open-source technologies such as Kafka, Spark, Iceberg, and Airflow.

A key part of our mission is to enable AI-centric automations that enhance the overall efficiency and intelligence of the platform. We are looking for passionate engineers who thrive on solving complex infrastructure challenges at scale, both on-premises and in the cloud. If you are dedicated to optimizing scalable, maintainable, and user-friendly systems, you will find compelling opportunities to make a significant impact at AiDP.

Description

The Service Reliability Engineer (SRE) role within AiDP Production Engineering is a dynamic position that blends strategic architectural design with hands‑on technical execution. As an SRE, you will be responsible for configuring, tuning, and ensuring the resilience of complex, multi‑tiered systems to achieve optimal application performance, stability, and availability. Our team manages critical data pipelines and applications across both bare‑metal and cloud computing platforms, delivering essential data processing for all of Apple’s key business functions.

We operate at an immense scale, handling exabytes of data, petabytes of memory, and tens of thousands of jobs to enable predictable and performance data analytics that power features and inform decisions across the company. If you are passionate about designing, building, and running data infrastructure that has a direct and significant impact on Apple’s global business operations, this is the ideal opportunity for you.

Responsibilities
  • Ability to understand the application requirements (Performance, Security, Scalability etc.) and assess the right services/topology on AWS, Baremetal & Kubernetes.
  • Build automation to enable self‑healing systems.
  • Build tools to monitor high performance & alert the low latency applications.
  • Ability to troubleshoot application specific, core network, system & performance issues.
  • Involvement in challenging and fast paced projects supporting Apple’s business by delivering innovative solutions.
  • Partner with engineering teams to prioritize and fix production defects.
  • Take knowledge transition from engineering teams for changes being rolled out in production.
  • Triage incidents based on the impact, devise and implement mitigation steps to unblock the business.
  • Conduct RCA, log defects and partner with engineering team for prioritization.
  • Support java based applications & Spark/Flink jobs on Baremetal, AWS & Kubernetes.
  • Share on‑call rotation with other team members to support apps and services in scope.
Minimum Qualifications
  • 4+ years experience in cloud‑native services, including ETL frameworks like Apache Spark, and Flink.
  • 4+ years experience in messaging systems (Kafka) and cloud infrastructure & services, AWS, GCP, Kubernetes.
  • 4+ years of experience in modern & distributed databases such as Snowflake, Cassandra, Single Store, and SAP HANA.
  • 4+ years of programming experience in Python or Java.
  • BS/MS in computer science or equivalent experience.
Preferred Qualifications
  • Solid understanding of system design, data structures, and incident management best practices.
  • Should be able to understand complex architectures and be comfortable working with multiple teams.
  • Observability tools (e.g: Prometheus, Grafana, Cloud Watch).
  • Ability to conduct performance analysis and troubleshoot large scale distributed systems.
  • Should be highly proactive with a keen focus on improving uptime/availability of our mission critical services.
  • Strong expertise in troubleshooting complex production issues.
  • Excellent problem solving, critical thinking, and communication skills.
  • Proven ability…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary