SRE Engineer
Job in
London, Greater London, W1B, England, UK
Listed on 2026-09-05
Listing for:
Gazelle Global Consulting Ltd
Full Time
position Listed on 2026-09-05
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Key Responsibilities:
Experience:
10+ years of hands-on experience in Site Reliability Engineering (SRE), Dev Ops, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration.
Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, paired with heavy production experience managing Kubernetes clusters.
Instrumentation & Coding: Proficiency in systems or scripting languages (e.g., Python, Go, Bash, or JavaScript) and experience instrumenting applications for APM.
Datadog Mastery: Deep understanding of Datadog's core pillars-Infrastructure, APM, Logs, Metrics, Synthetics, and Security Monitoring. Datadog Certifications are a strong plus.
Problem-Solving Mindset: Demonstrated ability to debug complex, distributed microservices architectures under high-pressure incident response scenarios.
Communication
Skills:
Excellent interpersonal and stakeholder management skills, with the ability to translate technical telemetry data into actionable business and engineering insights.
Essential skills/knowledge/experience:
Architecture & Implementation
Platform Ownership: Design, deploy, and manage Datadog agents, integrations, and custom metrics across multi-cloud (AWS/Azure/GCP) and containerized (Kubernetes, Docker) environments.
Observability Pipelines: Architect and scale high-throughput log processing, routing, and transformation systems using Datadog.
APM & Infrastructure Monitoring: Configure and optimize Application Performance Monitoring (APM), Distributed Tracing, Real User Monitoring (RUM), and Infrastructure metrics.
Governance & Best Practices
Standardization: Establish company-wide standards for dashboards, monitors, SLOs/SLIs, and alert routing (integrating with Pager Duty, Jira, Opsgenie, etc.).
Cost & Performance Optimization: Audit and optimize Datadog usage, index management, log retention policies, and custom metric volume to maximize ROI and control licensing costs.
Security & Compliance: Leverage Datadog Security products (CSPM, CWPP, Cloud SIEM, Container Security) to maintain compliance postures and mitigate runtime threats.
Collaboration & Enablement
Cross-Functional Mentorship: Act as the go-to escalation point and technical mentor for Dev Ops, SRE, and Software Engineering teams regarding troubleshooting and instrumentation.
Training & Documentation: Create internal documentation, runbooks, and training modules to elevate organizational proficiency in observability.
Vendor Management: Act as the primary technical point of contact for Datadog account teams,
Contract:
6 months+
Rates: Excellent
Location:
York
Interview: 2 stages, 1 technical + 1 assessment
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×