Remote ML Infrastructure Lead
York, North Yorkshire, YO90, England, UK
Listed on 2026-07-20
-
IT/Tech
Machine Learning/ ML Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AI Engineer (Applied/Software)
ML Infrastructure Lead
Role:
Senior ML Infrastructure Lead focused on building and scaling the technical foundations for production ML rid role across ML infrastructure, platform engineering and MLOps.
Reports to:
Chief Scientific Officer
Location:
WeWork Waterloo – Hybrid
Compensation:
Negotiable (Base) + 20% Company Performance Bonus + Share Options + iProov Benefits
We are looking for a highly capable and hands-on leader to design, implement and scale the systems, tooling, processes and standards that enable ML teams to train, deploy, monitor and improve models reliably, securely and at scale.
The role sits at the intersection of machine learning, software engineering, data, cloud infrastructure and platform reliability, bridging research and production. It suits someone who can think strategically about long-term platform capability while remaining technically hands-on to solve complex engineering and operational challenges.
Impact:
Lead the design and evolution of our ML platform, infrastructure and MLOps; build and maintain scalable, reliable, and secure systems for model training, testing, deployment, monitoring and lifecycle management; develop tooling to enable ML Engineers, Data Scientists and Researchers to work efficiently and ship models with confidence; define robust CI/CD workflows, model versioning, reproducibility, experimentation, feature management and release management;
own and improve production ML environments with strong standards for availability, performance, observability and resilience; monitor model and platform health, data quality, drift, latency, throughput and cost efficiency; build self-service tooling to reduce friction for ML teams; partner with ML, Data, Software and Platform Engineering to product ionise models and improve the end-to-end lifecycle; support scaling for training and inference workloads including high-throughput or compute-intensive use cases;
drive governance, security, compliance, auditability and operational rigor across the ML lifecycle; improve efficiency and cost-effectiveness of ML systems; mentor engineers and act as a technical leader; help define the roadmap for ML enablement.
What we would like to see from you
You will have experience in high-growth, fast-paced tech environments and are passionate about building and launching quality products with positive impact. You are an experienced product leader with a background in security (IAM) or enterprise SaaS, combining strategic vision with operational rigor to deliver usable, secure, and elegant technical solutions.
- Proven experience in a senior MLOps, ML Platform, ML Infrastructure, Platform Engineering or Machine Learning Systems role
- Strong hands-on background in software engineering and cloud infrastructure, with direct experience supporting production ML environments
- Experience building and operating systems that support the full ML lifecycle (experimentation, training, deployment, monitoring)
- Strong knowledge of Python and engineering practices (testing, automation, code quality)
- Strong experience with cloud platforms such as GCP
- Experience with Docker, Kubernetes and modern containerised deployment patterns
- Strong experience with CI/CD, infrastructure-as-code and workflow orchestration
- Experience with tools like Airflow or similar platforms
- Understanding of model observability, data quality, feature pipelines, lineage and reproducibility
- Experience designing scalable infrastructure for ML workloads (training, batch inference and real-time serving)
- Commitment to reliability, security, governance and operational excellence in production systems
- Ability to operate across both strategic and hands-on technical work
- Strong communication skills and ability to collaborate across engineering, product and data teams
Nice-to-haves
- Experience supporting computer vision, deep learning, LLM or compute-intensive ML workloads
- Experience with GPU infrastructure, distributed training or HPC environments
- Familiarity with feature stores, model registries and automated retraining pipelines
- Experience building internal developer platforms or self-service ML tooling
- Experience in regulated, high-security or…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: