Site Reliability Engineer Lead
Listed on 2026-08-24
-
IT/Tech
Systems Engineer, SRE/Site Reliability
Job Title
Required Qualifications
Strong experience in Site Reliability Engineering / Production Engineering
Hands-on expertise with:
IBM MQ (queue managers, clustering, channels, DLQ management)
Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
Large-scale distributed messaging systems and runtime management
Deep understanding of:
System reliability, scalability, and high availability design
Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering)
Incident management, root cause analysis, and problem management
Experience with:
Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms
Event and anomaly detection in high-volume systems
Strong scripting/automation skills:
Shell, Python, Power Shell
Experience managing Linux/Unix and Windows production environments
Knowledge of:
Event-driven architecture and messaging-based integration patterns
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).