Site Reliability Engineer III
Job in
Belfast, County Antrim, BT1, Northern Ireland, UK
Listed on 2026-08-29
Listing for:
CME- Group
Full Time
position Listed on 2026-08-29
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Job Description & How to Apply Below
Site Reliability Engineer (SRE) III – Platform Engineering & Systems Reliability
The Role:
CME Group is seeking a Site Reliability Engineer (SRE) III to engineer reliability for our Google Cloud (GCP) infrastructure, Middleware Platform Engineering team, and core technology foundations powering our Clearing, Risk, and derivatives applications. In this role, you will help build resilient, automated systems that combine ultra-low latency with high-concurrency performance, enabling CME's product teams to innovate safely will work alongside senior engineers, mentor junior colleagues, engage in the dynamic operation of production systems, and assist in driving our cloud transformation.
What You Will Do /
Key Responsibilities Middleware & Application Architecture:
Architect, operate, and support the migration of application platforms—including Messaging (Kafka, Red Panda, MQ, Pub/Sub), Service Discovery (Consul, Vault), and Data Distribution (SFTP/JScape)—to Google Cloud Platform. Manage cluster life cycles, data replication, RBAC, and workload placement.
Observability & Monitoring Fabric:
Design, scale, and maintain our observability backbone using tools like Open Telemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs, and SLOs to enable fast issue detection.
Incident Response & Operations:
Engage with urgency in live production incidents, take ownership of minor incidents, lead post-mortems, and ensure rapid system recovery.
Toil Reduction & Automation:
Actively identify operational toil and eliminate manual effort through code, automation, and systematic platform improvements.
Resiliency & Testing:
Contribute to disaster recovery (DR) strategies, continuous systems resiliency testing, and present reliability improvement suggestions to the Product backlog.
Collaboration & Leadership:
Lead technical discussions for assigned scope, present solution options, collaborate across functional teams, and mentor junior SRE colleagues.
What We're Looking For Engineering & Scripting Discipline:
Programming and scripting skills in high-level languages such as Python, Go, Java, or Bash to construct production-grade tooling.
Cloud Native & Systems Fundamentals:
Proficiency with Linux-based systems, distributed systems, containerization (Kubernetes/GKE), and public cloud platforms (GCP/GCE).Infrastructure as Code (IaC):
Understanding of modern CI/CD patterns and IaC tools such as Terraform, Ansible, or Kubernetes Config Connector (KCC).Networking & Protocols:
Knowledge of core systems and networking concepts (TCP/IP, UDP, HTTP, DNS, load balancing, and messaging protocols).AI & Agentic Engineering:
Forward-thinking approach to automation, leveraging Generative AI and Agents (e.g., Gemini) to optimize platform operations.
Analytical Problem-Solving:
Data-driven mindset to troubleshoot complex, non-linear system behaviors in a fast-paced, high-pressure trading ecosystem.
Communication & Adaptability:
Strategic communication skills to translate technical requirements for cross-functional teams, coupled with an eagerness to learn independently and collaboratively.
Preferred Qualifications / Desirable Observability Stack:
Hands-on experience with telemetry tools such as Open Telemetry, Splunk, Prometheus, and Grafana.
Agile Integration:
Comfort working within Agile frameworks and collaborative software development life cycles.
Certifications:
GCP Professional Cloud Architect, Certified Kubernetes Administrator (CKA), or Certified Kubernetes Application Developer (CKAD).Domain Expertise:
Any experience in Financial Markets or other highly regulated, ultra-low latency, high-concurrency environments would be highly beneficial although not essential," Why CME Group?
Global Significance:
Build technology that underpins the integrity of the world's leading derivatives marketplace.
Engineering Culture:
Flourish in a "code-first" environment that prioritizes systematic, automated solutions over manual intervention.
Professional Evolution:
Grow your SRE career within an organization actively transforming its approach to production engineering.
Competitive Package:
Enjoy…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×