Platform Engineer
Listed on 2026-08-06
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Job Summary
We are looking for a skilled Kubernetes Platform Engineer to help us stabilize, scale, and evolve our Azure Kubernetes Services (AKS) environment. This role works closely with application, infrastructure, and services teams to deliver, support, and maintain our Kubernetes fleet. Responsibilities span Developer Experience (Dev Ex), CI/CD tooling - including Git Hub Actions, Harness, and Flux CD - as well as observability platforms (primarily Dynatrace) that underpin our application estate.
This position is expected to be a thought leader in the Kubernetes and Dev Ops space, championing the elimination of technical debt, improving self-service developer tooling, and driving consistent, repeatable delivery practices across engineering teams.
- Manage, operate, and continuously improve Kubernetes clusters on Azure Kubernetes Service (AKS) across dev, staging, and production environments.
- Design and manage IaC (Terraform / Bicep) components to provision primary and supporting infrastructure in a repeatable, auditable manner.
- Optimize cluster performance, scalability, cost, and security posture; implement resource quotas, autoscaling (HPA/VPA/KEDA), and node pool strategies.
- Troubleshoot and resolve issues related to AKS, container orchestration, networking (CNI, Kong Ingress Controller, Kong Mesh), and storage.
- Maintain and enforce Kubernetes RBAC, Pod Security Standards, and network policies in alignment with security baselines.
- Implement, maintain, and improve CI/CD pipelines using Git Hub Actions as the primary pipeline platform - including reusable workflows, composite actions, and environment-based deployment gates.
- Administer and extend Harness for continuous delivery orchestration, including pipeline templates, approval workflows, canary/blue-green deployment strategies, and feature flags.
- Manage Git Ops-based deployments with Flux CD; maintain Helm releases, Kustomizations, and image update automation.
- Define and enforce branching strategies, release management processes, and promotion gates across environments.
- Integrate security scanning (container image scanning, SAST, secret detection) into CI pipelines to shift-left on vulnerabilities.
- Maintain artifact management practices using Git Hub Packages or Azure Container Registry (ACR).
- Act as the primary platform engineering point of contact for application development teams, enabling self-service onboarding to AKS and CI/CD tooling.
- Build and maintain internal developer tooling, golden-path templates, and paved-road patterns, including:
- Helm chart libraries
- Git Hub Actions workflow templates
- Harness pipeline templates
- Provide hands-on technical support and guidance to developers troubleshooting:
- Deployment failures
- Container issues
- Resource constraints
- Pipeline errors
- Run enablement sessions, write runbooks, and maintain internal documentation to reduce friction and support engineering velocity.
- Partner with application teams during architecture design and pre-production readiness reviews to ensure workloads are Kubernetes-native and production-ready.
- Champion developer feedback loops by collecting pain points and translating them into platform improvements.
- Monitor and manage Kubernetes resources and infrastructure using Dynatrace as the primary observability platform; build and maintain dashboards for cluster health, workload SLIs, and pipeline performance.
- Define and track SLOs/SLAs for platform components; lead incident response and post-mortem processes for platform-level outages.
- Leverage Dynatrace for:
- Full-stack APM
- Distributed tracing
- AI-assisted anomaly detection (Davis)
- Integrate Open Telemetry instrumentation for application-level observability.
- Maintain log ingestion and analysis pipelines within:
- Dynatrace (Log Monitoring v2)
- Azure Monitor Logs for centralized developer and operations visibility.
- Identify and address technical debt across the Kubernetes and Dev Ops toolchain; drive proactive remediation before it impacts reliability.
- Evaluate, recommend new tooling, and platform capabilities;…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).