×
Register Here to Apply for Jobs or Post Jobs. X

Cloud Operations Specialist

Job in Cape Town, 7561, South Africa
Listing for: Old Mutual South Africa
Full Time position
Listed on 2026-08-23
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Job Description & How to Apply Below

Let's Write Africa's Story Together!

Old Mutual is a firm believer in the African opportunity and our diverse talent reflects this.

Let's Write Africa's Story Together!

Old Mutual is a firm believer in the African opportunity and our diverse talent reflects this.

Job Description

The Cloud Operations Specialist is part of the team that is responsible for managing and maintaining Old Mutual's cloud infrastructure, ensuring high availability, performance, and security. The role is primarily operational in nature — focused on monitoring, incident response, system maintenance, patching, and the reliable day-to-day running of cloud platforms and shared services.

Cloud Operations Specialists work closely with Cloud Engineers, consuming and extending the automation tooling and infrastructure built by engineering teams. They are expected to write and maintain operational scripts to support their work, but their primary output is stable, well-operated cloud environments rather than the delivery of new features or platform capabilities.

A key and ongoing responsibility of this role is cloud cost management and Fin Ops practice. Cloud Operations Specialists are the primary owners of day-to-day spend visibility, waste identification, rightsizing, and cost reporting. Optimising cloud spend is not a periodic task — it is a continuous operational discipline embedded in the daily and weekly rhythm of the role.

This role additionally carries formal people management responsibility, leading the Cloud Operations team, managing team performance and development and acting as the senior escalation point for operational issues.

Infrastructure Operations & Availability
  • Manage and maintain cloud infrastructure to ensure high availability, performance, and security.
  • Assist in the deployment, configuration, and management of cloud resources in line with engineering-defined standards.
  • Perform regular system maintenance, updates, backups, and patching to keep cloud environments secure and up-to-date.
  • Manage infrastructure availability, resilience, and disaster recovery procedures.
Monitoring & Capacity Management
  • Monitor cloud systems and services to ensure optimal performance.
  • Implement and manage monitoring systems and monitoring reporting in line with engineering-defined requirements.
  • Monitor, report, and manage cloud platform and shared services capacity.
  • Optimize cloud resource usage and costs, providing recommendations for improvements.
  • Monitor cost and usage dashboards daily; identify anomalies, unexpected spend spikes, or untagged resources and elevate or remediate promptly.
Incident Management & Support
  • Respond to and resolve cloud-related incidents and service requests per SLA.
  • Participate in on-call rotations to provide 24/7 support for critical cloud operations.
  • Perform Tier-2 troubleshooting and elevate complex or recurring issues to Cloud Engineering for root cause analysis and remediation.
  • Collaborate with development and engineering teams to support cloud-based applications and services.
Fin Ops & Cloud Cost Optimisation
  • Contribute to cloud cost visibility tooling and dashboards (e.g. AWS Cost Explorer, Azure Cost Management, GCP Billing, Custom Implementations), ensuring spend is accurately tagged and attributable to teams, products, and environments.
  • Perform regular (at minimum once per sprint) reviews of cloud spend to identify waste, idle resources, oversized instances, unattached storage, and unused reservations.
  • Execute rightsizing recommendations generated by cloud-native tooling and engineering — resize, downgrade, or terminate resources within agreed operational parameters.
  • Consult on Reserved Instance / Savings Plan / Committed Use Discount coverage, flagging expiries and utilization gaps to Fin Ops and management.
  • Enforce and monitor resource tagging compliance; identify and remediate untagged or incorrectly tagged resources using automation tooling provided by engineering.
  • Consult and operate automated cost governance policies (budget alerts, anomaly detection rules, resource lifecycle schedules such as start/stop automation).
  • Identify and implement quick-win cost reduction opportunities (e.g. storage tier transitions, snapshot cleanup, log retention enforcement) within operational authority.
  • Contribute operational context and findings to quarterly Fin Ops reviews and cost deep-dives.
  • Collaborate with Cloud Engineering to raise architectural cost drivers that require design changes beyond operational remediation.
Automation & Operational Scripting
  • Operate and execute automation tooling and IaC workflows authored and maintained by Cloud Engineers.
  • Write and maintain operational scripts (Terraform, Python, Bash, or Power Shell) to support day-to-day maintenance, remediation, and reporting tasks.
  • Contribute improvements to runbooks and operational automation, working in partnership with engineering for any changes to production IaC or platform tooling.
Governance, Security & Compliance
  • Ensure compliance with security policies, organizational…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary