Sr. Platform Engineer, ML Infrastructure
Job in
Cambridge, Middlesex County, Massachusetts, 02140, USA
Listed on 2026-08-22
Listing for:
Insilico Search Partners
Full Time
position Listed on 2026-08-22
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Job Description & How to Apply Below
About the Company
Our client is a venture-backed biotech company applying AI to drug discovery, using a proprietary platform to identify novel drug targets and therapeutics from complex biological data.
Own and evolve the shared infrastructure — Kubernetes, Terraform, CI/CD, orchestration, storage, security, observability, and GPU systems — behind the company's ML and scientific workloads, across on-prem and cloud. You'll work closely with Machine Learning Engineers and researchers to keep training, evaluation, and inference reliable and scalable. Infrastructure-first role, with meaningful ML systems ownership (~60/40 split).
Key Responsibilities- Manage infrastructure as code (Terraform) and operate production Kubernetes clusters — networking, IAM/RBAC, security, resource allocation, reliability.
- Build and maintain CI/CD and Git Ops workflows (Git Hub Actions, CircleCI, Helm, Argo CD); secure and optimize container environments.
- Operate workflow orchestration (Kubeflow, Argo Workflows, Nextflow, or Seqera) and storage infrastructure for large scientific/ML datasets.
- Manage GPU-enabled infrastructure: provisioning, scheduling, drivers, utilization, troubleshooting.
- Establish observability (metrics, logs, dashboards, alerting) and document architecture/runbooks so operations don't depend on one person.
- Deploy and operate ML training/inference infrastructure (Anyscale, Ray, Vertex AI, Kubernetes); own shared tooling for experiment tracking and model/data versioning.
- Partner with ML Engineers to move model workloads onto shared, production-ready infrastructure.
- Lead cross-functional infrastructure initiatives and mentor engineers/scientists on the platform.
- 5+ years in production infrastructure, platform engineering, Dev Ops, SRE, or MLOps.
- Strong hands-on Kubernetes and infrastructure-as-code (Terraform) experience.
- Experience with major cloud platforms, CI/CD, containerization, and GPU/compute-intensive workloads.
- Workflow orchestration and Git Ops tooling (Kubeflow, Argo, Nextflow, Helm, Kueue).
- Familiarity with PyTorch and distributed ML execution; experience with model-serving platforms (Ray, Vertex AI).
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×