Devops & Platform Engineer
Listed on 2026-08-31
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Firm Overview
Lunate is a new Abu Dhabi-based, Partner-led, independent global alternative investment manager with more than 200 employees and $115 billion of assets under management. Lunate invests across the entire private markets spectrum including buyouts, growth equity, early and late-stage venture capital, private credit, real assets, and public equities and public credit. Lunate aims to be one of the world’s leading private markets solutions providers through SMAs and multi‑asset class funds, seeking to generate best‑in‑class risk‑adjusted returns for its clients.
Role OverviewLunate is seeking a Senior Lead Dev Ops / SRE Engineer to take ownership of how architecture is delivered into production across our Azure estate. This is a hands‑on technical role for someone who can decompose complex solution designs and land them as secure, reliable, production‑grade platforms.
The successful candidate will lead Azure Kubernetes Service (AKS) platform engineering, own Infrastructure as Code delivery using Terraform and Ansible, engineer Azure Dev Ops CI/CD pipelines, and establish the release management and Site Reliability Engineering practices that keep Lunate’s platforms fast, safe, and observable.
The role sits within the Technology Architecture function and works closely with cloud engineering, infrastructure engineering, security, and application teams.
Key ResponsibilitiesArchitecture Delivery
Translate reference architectures and solution designs into implementable, production‑ready platforms, covering security, networking, identity, and runtime considerations. Own end‑to‑end delivery of complex platform changes through design decomposition, implementation planning, engineering execution, and operational handover. Drive technical trade‑offs across reliability, security, performance, and maintainability.
AKS Platform Engineering and Operations
Design, implement, and operate AKS clusters including upgrades, scaling, node pool strategy, cluster hardening, workload isolation, and performance optimisation. Troubleshoot complex issues spanning Kubernetes, networking, identity integration, and underlying infrastructure, with a bias toward root‑cause resolution and automation.
Infrastructure as Code (Terraform and Ansible)
Build and maintain IaC frameworks using Terraform for provisioning and Ansible for configuration management. Create reusable modules, roles, and playbooks; enforce coding standards; manage drift; and integrate IaC into CI/CD workflows for controlled promotion across environments. Deliver end‑to‑end automation from infrastructure provisioning through post‑provision configuration and hardening.
CI/CD Engineering (Azure Dev Ops)
Design and implement multi‑stage Azure Dev Ops pipelines, deployment strategies, approval gates, artifact management, and rollback patterns. Standardise pipelines through reusable templates and documented patterns to improve quality, repeatability, and auditability.
Release Management
Own the release management operating model, including release planning, change control alignment, release communications, promotion sequencing, verification, and rollback readiness. Use reliability signals and operational readiness criteria to govern release confidence and reduce deployment risk.
Automation and Scripting
Develop automation and tooling in Python, Bash, and Power Shell to reduce operational toil and improve repeatability across build, deploy, monitoring, and incident response workflows.
Linux and Windows Server Engineering
Operate and troubleshoot both Linux and Windows Server environments supporting application and platform workloads, applying consistent configuration and hardening practices across both estates.
Site Reliability Engineering
Define and ope rationalise SLIs, SLOs, and error budget policies to balance delivery velocity with service reliability. Lead incident response, establish runbooks, and drive blameless postmortems with measurable actions to reduce recurrence.
Technical Systems and ToolingCloud and Orchestration
Microsoft Azure, Azure Kubernetes Service (AKS), Azure networking and identity services.
Containerisation
Docker (image build, multi‑stage builds,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).