Storage Platform Engineer
Listed on 2026-06-26
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Title: Storage Platform Engineer
Contract Length: 6 Months ASAP Start
Day Rate: Inside IR35
Hybrid: 2 days on site (Cambridge)
Job DescriptionThe Storage Platform Engineer role sits within the Client's Global Storage team, supporting large-scale storage platforms used by engineering and HPC workloads across on-premises and cloud environments.
This role focuses on making storage platforms reliable, observable, and easier to operate. It includes reducing manual work through automation, building practical tooling, and helping teams use storage services effectively at scale.
Working with colleagues across multiple regions, the role contributes to resolving issues, addressing root causes, and maintaining stable, well-performing systems that support the Client's technology development.
Responsibilities- Maintain the reliability, availability, and performance of storage platforms used by engineering teams.
- Contribute to incident response, investigation, and problem resolution.
- Apply service reliability measures such as SLOs and SLIs where appropriate.
- Build and maintain infrastructure using Terraform and Ansible.
- Develop automation and Python-based tools to support operations and system insight.
- Use AI-based tooling to assist with monitoring, anomaly detection, and analysis.
- Develop simple agent-based workflows to support operational decision-making.
- Enhance monitoring and alerting to provide clear visibility of system behaviour.
- Work with engineering and security teams to maintain secure and well-managed systems.
- Maintain accurate documentation and share knowledge across the team.
- Administering Net App storage platforms in production environments.
- Experience with Linux operating system and supporting large-scale Linux HPC environments, including file and object storage.
- Experience with Infrastructure as Code tools such as Terraform or configuration tools such as Ansible.
- Ability to develop automation or tooling using a programming language such as Python.
- Experience supporting reliable and scalable systems in an operational environment.
- Familiarity with AWS, GCP, or Azure.
- Exposure to CI/CD or Git-based workflows.
- Experience using or integrating AI/ML or agent-based tooling in operations.
- Understanding of identity, access control, and security practices.
- Experience with Open Stack storage services such as Cinder or Manila.
- Awareness of service management approaches (e.g. ITIL).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).