×
Register Here to Apply for Jobs or Post Jobs. X

Senior DevOps​/SRE Cloud Engineer

Job in Cape Town, 7561, South Africa
Listing for: Hire Resolve
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below

A company that develops, owns, and operates utility-scale power generation and energy storage plants across the African continent is seeking a Senior Dev Ops / SRE Cloud Engineer who will own the company's Azure-based Kubernetes platform end-to-end - remote or hybrid based on location .

Responsibilities:
  • Infrastructure as Code: Provision and operate production Kubernetes (AKS) and Azure infrastructure using modular Terraform.

  • CI/CD Automation: Maintain Git Hub Actions pipelines featuring quality gates, security checks, and automated deployments.

  • Observability & SRE: Own monitoring, alerting, SLOs, capacity planning, and incident response across environments.

  • Security & Access: Manage platform secrets, network controls (VNets, private endpoints), and Entra  governance.

  • Developer Enablement: Partner with engineering teams to resolve operational bottlenecks and drive reliability standards.

Minimum Requirements:
  • Experience: 5+ years in Dev Ops/SRE/Platform Engineering (including 2-3 years operating K8s in production and 2+ years on Azure).

  • Education: Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.

  • Certifications: CKA, CKAD, or Azure Solutions Architect / Dev Ops Engineer certifications are advantageous.

  • Soft Skills: Strong problem-solving ability, clear technical communication, and the capacity to operate autonomously within a hybrid/remote team.

  • Azure Platform: Production experience with AKS, ADLS Gen2, Key Vault, Entra  (Workload/Managed Identities), VNets/Private Endpoints, and Service Bus (or equivalent broker).

  • Production Kubernetes: Advanced operational depth in Helm chart authoring, operator deployment, node-pool sizing, pod troubleshooting, and cluster upgrades.

  • Terraform: Proven ability to author modular IaC, manage state safely, and maintain strict plan/apply disciplines.

  • Containerization & CI/CD: Hands-on Docker (multi-arch builds, optimization) and Git Hub Actions pipeline development with required status checks.

  • Observability: Practical experience configuring Prometheus + Grafana, structured logging, and SLO-based alerting.

  • Linux & Automation: Strong Bash scripting with clean, code-maintained operational tooling.

Preferred Experience
  • Data Platforms: Apache Spark on K8s (Spark Operator/Connect), Jupyter Hub, Delta Lake, Trino, Hive Metastore, or Apache Ranger.

  • Advanced Telemetry: Open Telemetry (SDKs/Collector topology) and Open Lineage/Marquez integration.

  • Languages & Databases: Intermediate Python (infrastructure testing/pytest) and basic DBA management for SQL Server and PostgreSQL.

  • Compliance & Isolation: Multi-tenant architecture design, dependency auditing, image provenance, and POPIA/GDPR/ISO 27001 controls.

Benefits:
  • Competitive salary based on experience (salary can potentially be more based on experience/skills)

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary