×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer - Ceph Storage

Job in Austin, Travis County, Texas, 78701, USA
Listing for: GoDaddy
Full Time position
Listed on 2026-07-10
Job specializations:
  • IT/Tech
    Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Job Description & How to Apply Below

Senior Site Reliability Engineer

GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience.

At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering.

As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide.

What You'll Get to Do…

  • Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage workloads (RADOS, RGW, RBD, CephFS).
  • Diagnose and resolve complex distributed systems issues including recovery/backfill events, OSD instability, PG imbalance, storage latency, and RGW performance bottlenecks.
  • Design and build automation using Python, Shell, Salt Stack, and Ansible to reduce operational toil and improve platform resilience at scale.
  • Define and improve observability for the storage platform through SLIs, SLOs, PromQL, LogQL, Grafana dashboards, and proactive alerting strategies.
  • Lead storage lifecycle initiatives including cluster expansions, hardware refreshes, software upgrades, Open Stack integrations, and large-scale migration projects.

Your Experience Should Include…

  • 5+ years operating large-scale Linux infrastructure with significant experience supporting distributed storage systems in production environments.
  • 2+ years hands-on Ceph administration including OSD, MON, MDS, and RGW operations, CRUSH map management, pool design, placement groups, and performance troubleshooting.
  • Strong understanding of Linux internals, storage architecture, networking, file systems, block devices, and performance analysis under production workloads.
  • Experience developing operational automation and tooling with Python, Shell, and configuration management platforms such as Ansible, Salt Stack, Puppet, or Chef.
  • Proven incident response expertise, including production troubleshooting, root-cause analysis, postmortem creation, and implementation of durable corrective actions.

You Might Also Have…

  • Advanced Ceph expertise across RGW Multisite, CephFS, RBD Mirroring, erasure coding, and large-scale cluster design.
  • Experience integrating storage platforms with Open Stack (Cinder, Nova, Neutron, Swift) or Kubernetes storage technologies (CSI, Storage Classes, Rook).
  • Familiarity with modern observability platforms including Prometheus, Mimir, Loki, Grafana, and enterprise monitoring architectures.
  • Experience planning and executing petabyte-scale storage migrations, capacity expansion programs, and storage hardware lifecycle management.
  • Contributions to open-source infrastructure or storage projects and a passion for advancing distributed systems technologies.

We encourage you to apply even if your experience or skillset doesn't align perfectly with every requirement. We value a wide range of backgrounds and transferable skills, and we are excited to support learning and growth.

GoDaddy is proud to be an equal opportunity employer. GoDaddy will consider for employment qualified applicants with criminal histories in a manner consistent with local and federal requirements. Refer to…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary