Staff Site Reliability Engineer
Listed on 2026-08-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Cybersecurity
Staff Site Reliability Engineer
Santa Clara, California, United States
IonQ, Inc. is the world's leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ's newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting-edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense.
In 2025, the company achieved 99.99% two-qubit gate fidelity, setting a world record in quantum computing performance. Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Canada, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space.
IonQ is making quantum platforms more accessible and impactful than ever before.
We are seeking a Staff Site Reliability Engineer. As Staff SRE Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing.
Responsibilities:
- Production reliability: own service-level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions.
- Engineer observability: design and operate the observability stack so production services are fully instrumented and define the standards platform and application teams follow.
- Govern SLOs and error budgets: define and manage service-level objectives, run regular reviews with service owners, and drive corrective action when services consume error budgets unsafely.
- Drive resilience: design and execute chaos experiments and validate that failure modes are covered by tested safeguards.
- Lead incident response: define the incident process and serve as incident commander for the highest-severity incidents, including security incidents within the coverage window.
- Run on-call and escalation: establish and manage rotations and escalation paths that provide continuous coverage with clean follow-the-sun handoffs.
- Disaster recovery: own disaster-recovery testing and failover validation against defined recovery objectives and turn exercise findings into architectural and operational improvements.
- Cloud security posture: co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring with Dev Sec Ops .
- Data, streaming, and AI Ops: own reliability of stateful and streaming services, capacity planning and rightsizing, and autonomous agents for triage, predictive alerting, remediation, and self-healing.
- Scale the team and broaden impact: mentor engineers at different seniority levels, set standards adopted across teams, and align Architecture, Dev Sec Ops , Cloud Operations, and Product Development behind a shared reliability roadmap.
Requirements:
- 7+ years of production engineering experience with recent hands-on reliability work.
- Hands-on, recent experience operating large-scale, fault-tolerant production systems on AWS or GCP.
- Observability ownership: have instrumented production systems and governed service-level objectives and error budgets, not only installed dashboards.
- Resilience practice: have designed and executed failure experiments or disaster-recovery exercises with real failover validation.
- Incident command: have personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix.
- Demonstrated ownership of reliability outcomes with measurable results, such as availability, mean time to recovery, and error-budget adherence.
- Evidence of multi-team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team.
Preferred Qualifications:
- Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments.
- Strong experience prioritizing risk using identity, workload, and exposure-path context to focus remediation on issues that materially increase attack likelihood and operational impact.
- Experience with autonomous remediation and self-healing workflows powered by AIOps, including Amazon Bedrock Agent…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).