×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Platform Engineer, ML Infrastructure

Job in Cambridge, Middlesex County, Massachusetts, 02140, USA
Listing for: Insilico Search Partners
Full Time position
Listed on 2026-08-22
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 140000 - 210000 USD Yearly USD 140000.00 210000.00 YEAR
Job Description & How to Apply Below

About the Company

Our client is a venture-backed biotech company applying AI to drug discovery, using a proprietary platform to identify novel drug targets and therapeutics from complex biological data.

Own and evolve the shared infrastructure — Kubernetes, Terraform, CI/CD, orchestration, storage, security, observability, and GPU systems — behind the company's ML and scientific workloads, across on-prem and cloud. You'll work closely with Machine Learning Engineers and researchers to keep training, evaluation, and inference reliable and scalable. Infrastructure-first role, with meaningful ML systems ownership (~60/40 split).

Key Responsibilities
  • Manage infrastructure as code (Terraform) and operate production Kubernetes clusters — networking, IAM/RBAC, security, resource allocation, reliability.
  • Build and maintain CI/CD and Git Ops workflows (Git Hub Actions, CircleCI, Helm, Argo CD); secure and optimize container environments.
  • Operate workflow orchestration (Kubeflow, Argo Workflows, Nextflow, or Seqera) and storage infrastructure for large scientific/ML datasets.
  • Manage GPU-enabled infrastructure: provisioning, scheduling, drivers, utilization, troubleshooting.
  • Establish observability (metrics, logs, dashboards, alerting) and document architecture/runbooks so operations don't depend on one person.
  • Deploy and operate ML training/inference infrastructure (Anyscale, Ray, Vertex AI, Kubernetes); own shared tooling for experiment tracking and model/data versioning.
  • Partner with ML Engineers to move model workloads onto shared, production-ready infrastructure.
  • Lead cross-functional infrastructure initiatives and mentor engineers/scientists on the platform.
Qualifications
  • 5+ years in production infrastructure, platform engineering, Dev Ops, SRE, or MLOps.
  • Strong hands-on Kubernetes and infrastructure-as-code (Terraform) experience.
  • Experience with major cloud platforms, CI/CD, containerization, and GPU/compute-intensive workloads.
  • Workflow orchestration and Git Ops tooling (Kubeflow, Argo, Nextflow, Helm, Kueue).
  • Familiarity with PyTorch and distributed ML execution; experience with model-serving platforms (Ray, Vertex AI).
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary