More jobs:
IT Platform Engineer - HPC & Linux
Job in
Milton Keynes, Buckinghamshire, MK1, England, UK
Listed on 2026-07-10
Listing for:
United States Digital Space LLC
Full Time
position Listed on 2026-07-10
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Infrastructure
Job Description & How to Apply Below
Purpose
Design, build, operate and continuously improve secure, scalable and reliable platform services supporting engineering, simulation and HPC workloads across hybrid cloud and on-premises environments. The role combines platform engineering, automation, infrastructure operations and developer enablement, delivering self-service capabilities and platform products that improve the productivity, reliability and efficiency of engineering teams.
Accountabilities- Design, implement and maintain innovative platforms end-to-end for the Technology Campus
- Design and develop self-service platform capabilities that enable engineering teams to provision and consume infrastructure, compute, storage and application services consistently and securely.
- Define and maintain service reliability objectives, capacity plans and operational metrics to ensure platform availability, performance and scalability.
- Provide advanced technical support and manage problem escalations from users communicating well and recording in ticketing management system.
- Partner with software engineering teams to understand application requirements, improve developer experience, support cloud-native delivery practices and ensure platform services effectively meet the needs of engineering workloads.
- Own on-premises and cloud-hosted platform services, proactively identifying reliability, performance and scalability improvements and leading the design and implementation of appropriate solutions.
- Drive automation and Infrastructure as Code practices to improve reliability, consistency, and operational efficiency of platform services
- Contribute to platform governance forums, providing technical input, sharing knowledge, and supporting alignment across infrastructure and engineering teams
Enable efficient and reliable consumption of platform services by internal teams, with a focus on usability, repeatability, and reduced operational friction
Additional Accountabilities- Participate in wider team projects, change management and take an active role in reviewing architecture of new solutions across Platforms infrastructure and HPC
- Contribute to platform architecture, technology roadmaps and lifecycle management decisions to ensure services remain fit for purpose and aligned with future business requirements.
- Maintain accurate platform documentation, asset records, and operational runbooks to support effective operation and knowledge sharing
- Ensure platform solutions comply with organisational security standards, policies, and regulatory requirements
- Work closely with vendors to leverage their expertise and solutions, making use of technical partnerships to improve performance, inform future technical direction, execute proof of concepts and to implement new technology
- Expert knowledge administering Linux/Unix systems
- Experience in scripting and programming e.g. Python, Bash, Go
- Experience implementing CI/CD and Git Ops workflows using tools such as Git Hub Actions, Git Lab, ArgoCD or Flux.
- Working knowledge of platform security principles including identity management, secrets management, vulnerability remediation, hardening and least-privilege access controls.
- Network knowledge of Infini Band, MPI and Ethernet concepts
- Working knowledge of platform security principles.
- Experience with configuration management and infrastructure as code tools (Ansible, Git, Vault)
- Knowledge of container orchestration tools, virtualisation and observability stacks e.g. Kubernetes, Grafana, Kafka, Docker, Open Stack, OLVM.
- Experience implementing platform and application observability solutions including monitoring, logging, tracing, alerting and telemetry.
- Awareness of secure software development and Dev Sec Ops principles.
- Experience working within Agile and Dev Ops delivery models.
- Knowledge of software delivery platforms such as Git Hub, Git Lab, Azure Dev Ops or equivalent.
- Experience designing, deploying and operating infrastructure services within public cloud platforms.
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×