×
Register Here to Apply for Jobs or Post Jobs. X

Sr Site Reliability Engineer

Job in Austin, Travis County, Texas, 78716, USA
Listing for: News Corporation
Full Time position
Listed on 2026-06-26
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, AWS
Salary/Wage Range or Industry Benchmark: 130000 - 160000 USD Yearly USD 130000.00 160000.00 YEAR
Job Description & How to Apply Below

Senior Site Reliability Engineer

This role contributes to the reliability, observability, and operational excellence of our platform infrastructure serving millions of users. As a Senior SRE, you will be a strong technical contributor who implements best practices, solves complex problems, and enables our 600+ engineers to deliver exceptional customer experiences.

What You'll Do:
Platform Reliability & Infrastructure
  • Implement and maintain highly available AWS infrastructure including EKS clusters, Fargate (ECS), and multi-region architectures
  • Support reliability of critical services:
    Skyway (CI/CD), Frontdoor (Tyk), Pantheon (Apollo GraphQL), and supporting infrastructure
  • Monitor SLIs, SLOs, and error budgets for Tier 1/2/3 systems; participate in architectural reviews for reliability and cost-efficiency
  • Implement reliability patterns including circuit breakers, graceful degradation, and automated failover
Observability & Cost Optimization
  • Implement observability solutions using New Relic for APM, distributed tracing, metrics, and logging for rapid troubleshooting
  • Build dashboards and alerts that reduce MTTD and MTTR; contribute to observability standards across teams
  • Identify infrastructure cost optimization opportunities and implement Fin Ops practices including rightsizing and resource lifecycle management
  • Support cost-conscious architecture decisions and CI/CD spend optimization (CircleCI, Argo CD)
Chaos Engineering & Incident Response
  • Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing
  • Participate in game day exercises and disaster recovery simulations; create runbooks and automation for resilience
  • Participate in on-call rotation for critical systems; conduct post-incident reviews and implement improvements
  • Support incident response processes and contribute to System Health Scorecard
Technical Contribution
  • Contribute as a strong technical individual contributor to the Operations Excellence team
  • Collaborate with Platform Engineering, Quality Engineering, and product teams on reliability initiatives
  • Support security initiatives including AWS Secrets Manager migration and compliance requirements (SOC 2, PCI, GDPR)
  • Contribute to Developer Experience metrics and platform adoption goals
  • May provide technical guidance to junior team members
What You'll Bring:
  • 5+ years in Site Reliability Engineering, Dev Ops, or Infrastructure Engineering with demonstrated success improving system reliability
  • Bachelor’s degree or equivalent experience
  • 3+ years hands‑on experience with AWS (EKS, EC2, RDS, S3, Cloud Watch, IAM) and Kubernetes including cluster management
  • Proficient programming skills (Python, Go, or Java) with infrastructure automation and Infrastructure as Code experience (Terraform, Cloud Formation)
  • Production experience with observability tools (New Relic, Datadog, Prometheus, Grafana, Splunk) and distributed systems
  • Experience with CI/CD platforms and Git Ops workflows (CircleCI, Argo CD, Jenkins); on-call rotation and incident response
  • Preferred:
    Exposure to chaos engineering tools, API Gateway technologies (Tyk/Kong), GraphQL federation (Apollo), cost optimization initiatives, Fin Ops principles
Technical Skills
  • Cloud &

    Infrastructure: AWS (EKS, Fargate, Lambda, VPC, Route
    53, Cloud Front), Kubernetes, Docker, Istio Service Mesh
  • CI/CD & Git Ops:
    Argo CD, CircleCI, Jenkins, Git Hub Actions
  • Observability:
    New Relic - APM, distributed tracing, metrics & logging;
    Splunk - logging
  • IaC & Automation:
    Terraform, Cloud Formation, Helm, Kustomize, Python/Go/Bash
  • Platform Services:
    Tyk Gateway, Apollo GraphQL, AWS Secrets Manager, Vault
  • Incident Management:
    Ops Genie, Pager Duty, Service Now
Professional Qualities
  • Strong communication skills with ability to explain technical concepts to diverse audiences
  • Collaborative approach working across engineering, product, and business teams
  • Self-motivated with ability to solve complex problems within established practices and policies
  • Data-driven decision making with customer-centric approach and empathy for developer experience
How We Reward You:
  • Inclusive and Competitive medical, Rx, dental, and vision coverage
  • Family forming…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary