×
Register Here to Apply for Jobs or Post Jobs. X

Senior DevOps & Site Reliability Engineer

Job in Sandton, 2172, South Africa
Listing for: Datonomy Solutions (Pty) Ltd
Full Time position
Listed on 2026-08-30
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Job Description & How to Apply Below

We are looking for an experienced
Senior Dev Ops & Site Reliability Engineer (SRE)to design, build and operate highly available, secure, scalable and automated enterprise technology platforms.

This is a senior hands-on engineering role spanning
Dev Ops, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and Dev Sec Ops .

The successful candidate will work across engineering and delivery teams to improve
platform reliability, deployment velocity, resilience, automation, operational efficiency and production performance
, while supporting mission-critical enterprise applications.

Key Responsibilities Dev Ops & Platform Engineering
  • Design, build and maintain
    cloud-native infrastructure and platform services
    .

  • Develop and maintain
    Infrastructure as Code (IaC)solutions.

  • Automate infrastructure provisioning, configuration and operational processes.

  • Build reusable engineering tools, deployment templates and platform components.

  • Establish and standardise platform engineering practices across multiple delivery teams.

  • Identify opportunities to reduce manual intervention and increase engineering automation.

CI/CD & Release Automation
  • Design, implement and maintain enterprise-grade
    CI/CD pipelines
    for application and infrastructure deployments.

  • Implement automated testing, security scanning, code-quality controls and release automation.

  • Enable automated deployments, rollback and recovery processes.

  • Improve deployment frequency while reducing change and deployment risk.

  • Continuously optimise software delivery and release-management processes.

Site Reliability Engineering
  • Implement and mature
    Site Reliability Engineering practices
    across production environments.

  • Define, monitor and manage
    Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs).

  • Improve application and platform
    availability, scalability, resilience and performance
    .

  • Lead production incident response, troubleshooting, problem management and
    Root Cause Analysis (RCA).

  • Drive proactive reliability improvements and reduction of technical debt.

  • Improve
    Mean Time to Detect (MTTD)and
    Mean Time to Recover (MTTR).

Azure Cloud Engineering
  • Design, implement and operate enterprise
    Microsoft Azure
    environments.

  • Work extensively with technologies such as:

    • Azure Kubernetes Service (
      AKS
      )

    • Azure App Services

    • Azure Networking

    • Azure Monitor

    • Azure Storage

    • Azure Identity Services

  • Design highly available and disaster-recovery-capable environments.

  • Optimise cloud environments for
    performance, resilience, security and cost
    .

  • Support hybrid-cloud and multi-cloud environments where required.

Containers & Kubernetes
  • Build, deploy and support containerised applications using
    Docker
    and
    Kubernetes
    .

  • Manage Kubernetes environments, particularly
    Azure Kubernetes Service (AKS).

  • Develop and maintain deployment configurations using
    Helm
    .

  • Support container-platform reliability, scalability and operational performance.

  • Open Shift experience would be advantageous.

Infrastructure as Code & Automation

Hands-on experience with technologies such as:

  • Terraform

  • Bicep

  • ARM Templates

  • Ansible

Candidates should be comfortable using Infrastructure as Code to build repeatable, scalable and governed enterprise infrastructure.

Monitoring & Observability
  • Implement comprehensive
    logging, monitoring, metrics, tracing and alerting
    .

  • Build operational dashboards and platform insights.

  • Establish enterprise observability standards.

  • Implement proactive and predictive monitoring.

  • Use observability information to improve application and infrastructure reliability.

Relevant technologies may include:

  • Dynatrace

  • Grafana

  • Prometheus

  • Elastic Stack / ELK

  • Splunk

  • Azure Monitor

  • Open Telemetry

Dev Sec Ops  & Security
  • Embed
    Dev Sec Ops practices throughout the software-delivery lifecycle.

  • Integrate security scanning and controls into CI/CD pipelines.

  • Support vulnerability identification, remediation and risk reduction.

  • Ensure cloud and platform environments comply with enterprise security and regulatory requirements.

  • Work closely with information-security teams to continuously improve platform security.

Technical Leadership
  • Provide technical leadership across Dev Ops, Cloud, Platform and SRE…

Position Requirements
10+ Years work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary