Site Reliability Engineer
Listed on 2026-07-01
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Support, SRE/Site Reliability
Staff Site Reliability Engineer
We're looking for a Staff Site Reliability Engineer to join our team, focusing on the core systems that power global financial markets. This isn't just about keeping the lights on; it's about pioneering the future of financial technology. As a member of our Clearing department, you'll be on the front lines, ensuring the integrity and performance of mission-critical systems that facilitate billions of dollars in daily transactions.
If you're a builder at heart, driven by a passion for creating ultra-reliable and resilient systems, you'll thrive here.
This is a hybrid role. You must be in our office 2+ days a week.
What You'll Get- A supportive environment fostering career progression, continuous learning, and an inclusive culture.
- Broad exposure to CME's diverse products, asset classes, and cross-functional teams.
- A competitive salary and comprehensive benefits package.
As a Staff Site Reliability Engineer, you'll be a visionary builder of our resilient infrastructure. You'll move beyond conventional operations to apply software engineering principles to every facet of our clearing systems.
- Pioneer solutions to guarantee the reliability, performance, and availability of our CME clearing and risk systems, where every millisecond and every transaction counts.
- Architect and implement cutting-edge solutions for application resiliency and fault tolerance.
- Drive automation and continuous improvement across the entire system lifecycle, eliminating manual toil and enhancing operational excellence.
- Integrate SRE principles directly into the software development lifecycle, embedding reliability from day one.
- Collaborate with cross-functional development and platform teams, providing expert-level guidance to deploy and maintain critical applications.
- Innovate and lead efforts to prevent incidents, enhance operational processes, and automate solutions at a global scale.
- Spearhead the adoption of observability and performance testing, guiding teams to a "build with SRE mindset" culture.
- Own the end-to-end operational integrity of products, understanding and contributing to the bigger picture of the organization.
- A strong academic background:
Bachelor's degree in Engineering, Computer Science, Information Technology, or a related field is strongly preferred. - Cloud expertise:
Hands-on experience deploying and operating applications using IaaS and PaaS on major cloud providers, preferably Google Cloud Services. - Coding fluency:
Proficiency in one or more of the following languages:
Java, Python, Bash, or Go. Typescript and/or Rust are a significant plus. - Infrastructure as Code (IaC) mastery:
Experience with tools such as GKE, Terraform, Cloud Formation, and Chef. - Proven reliability engineering skills:
Deep knowledge of SRE and security best practices, with a track record of implementing them into workflows. A solid understanding of performance testing tools is essential, along with the ability to help teams resolve complex performance issues. - Automation prowess:
Demonstrated experience with automation, CI/CD, orchestration, and configuration management. - Observability knowledge:
Familiarity with logging and observability platforms such as Open Telemetry and Prometheus. - A security-first mindset:
Strong understanding of security and compliance frameworks. - Problem-solving abilities:
Excellent written and verbal communication skills, with the ability to convey complex technical concepts clearly to both technical and non-technical audiences. - Strong collaboration skills:
An agile team player who is self-motivated and can work with minimal supervision while juggling multiple concurrent projects. - A passion for innovation: A continuous desire to learn and stay up-to-date with the latest technologies and industry trends.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).