Cloud Operations Specialist
Listed on 2026-08-23
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Let's Write Africa's Story Together!
Old Mutual is a firm believer in the African opportunity and our diverse talent reflects this.
Let's Write Africa's Story Together!Old Mutual is a firm believer in the African opportunity and our diverse talent reflects this.
Job DescriptionThe Cloud Operations Specialist is part of the team that is responsible for managing and maintaining Old Mutual's cloud infrastructure, ensuring high availability, performance, and security. The role is primarily operational in nature — focused on monitoring, incident response, system maintenance, patching, and the reliable day-to-day running of cloud platforms and shared services.
Cloud Operations Specialists work closely with Cloud Engineers, consuming and extending the automation tooling and infrastructure built by engineering teams. They are expected to write and maintain operational scripts to support their work, but their primary output is stable, well-operated cloud environments rather than the delivery of new features or platform capabilities.
A key and ongoing responsibility of this role is cloud cost management and Fin Ops practice. Cloud Operations Specialists are the primary owners of day-to-day spend visibility, waste identification, rightsizing, and cost reporting. Optimising cloud spend is not a periodic task — it is a continuous operational discipline embedded in the daily and weekly rhythm of the role.
This role additionally carries formal people management responsibility, leading the Cloud Operations team, managing team performance and development and acting as the senior escalation point for operational issues.
Infrastructure Operations & Availability- Manage and maintain cloud infrastructure to ensure high availability, performance, and security.
- Assist in the deployment, configuration, and management of cloud resources in line with engineering-defined standards.
- Perform regular system maintenance, updates, backups, and patching to keep cloud environments secure and up-to-date.
- Manage infrastructure availability, resilience, and disaster recovery procedures.
- Monitor cloud systems and services to ensure optimal performance.
- Implement and manage monitoring systems and monitoring reporting in line with engineering-defined requirements.
- Monitor, report, and manage cloud platform and shared services capacity.
- Optimize cloud resource usage and costs, providing recommendations for improvements.
- Monitor cost and usage dashboards daily; identify anomalies, unexpected spend spikes, or untagged resources and elevate or remediate promptly.
- Respond to and resolve cloud-related incidents and service requests per SLA.
- Participate in on-call rotations to provide 24/7 support for critical cloud operations.
- Perform Tier-2 troubleshooting and elevate complex or recurring issues to Cloud Engineering for root cause analysis and remediation.
- Collaborate with development and engineering teams to support cloud-based applications and services.
- Contribute to cloud cost visibility tooling and dashboards (e.g. AWS Cost Explorer, Azure Cost Management, GCP Billing, Custom Implementations), ensuring spend is accurately tagged and attributable to teams, products, and environments.
- Perform regular (at minimum once per sprint) reviews of cloud spend to identify waste, idle resources, oversized instances, unattached storage, and unused reservations.
- Execute rightsizing recommendations generated by cloud-native tooling and engineering — resize, downgrade, or terminate resources within agreed operational parameters.
- Consult on Reserved Instance / Savings Plan / Committed Use Discount coverage, flagging expiries and utilization gaps to Fin Ops and management.
- Enforce and monitor resource tagging compliance; identify and remediate untagged or incorrectly tagged resources using automation tooling provided by engineering.
- Consult and operate automated cost governance policies (budget alerts, anomaly detection rules, resource lifecycle schedules such as start/stop automation).
- Identify and implement quick-win cost reduction opportunities (e.g. storage tier transitions, snapshot cleanup, log retention enforcement) within operational authority.
- Contribute operational context and findings to quarterly Fin Ops reviews and cost deep-dives.
- Collaborate with Cloud Engineering to raise architectural cost drivers that require design changes beyond operational remediation.
- Operate and execute automation tooling and IaC workflows authored and maintained by Cloud Engineers.
- Write and maintain operational scripts (Terraform, Python, Bash, or Power Shell) to support day-to-day maintenance, remediation, and reporting tasks.
- Contribute improvements to runbooks and operational automation, working in partnership with engineering for any changes to production IaC or platform tooling.
- Ensure compliance with security policies, organizational…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: