ML Platform Engineer
Listed on 2026-10-05
-
IT/Tech
Machine Learning/ ML Engineer, Data Engineering
Datavant is the data collaboration platform trusted for healthcare. Guided by our mission to make the world’s health data secure, accessible and actionable, we provide critical data solutions for organizations across the healthcare ecosystem - including providers, health plans, researchers, and life sciences companies. From fulfilling a single patient’s request for their medical records to powering the AI revolution in healthcare, Datavanters are building the future of how data is connected and used to improve health.
By joining Datavant today, you’re stepping onto a driven and highly collaborative team that is passionate about creating transformative change in healthcare.
What We’re Looking ForWe are looking for a Staff ML Platform Engineer to join our Data & ML Platform organization on the ML Platform team. We own the paved road that takes a model from notebook to production without rebuilding it each time: training and inference across Sage Maker and Databricks, self-service LLM-endpoint serving, and the pipelines behind Datavant’s clinical-AI products. Data Science owns the models and their quality;
we own the pipelines, serving infrastructure, and operational guardrails that let those models run safely against regulated health data.
As a Staff ML Platform Engineer, you will be a technical leader across the ML Platform: setting direction for the paved road, owning the hardest architectural problems, and moving the team from bespoke plumbing toward a coherent platform that Data Science teams can self-serve. AI fluency is a baseline expectation. You should already be using Claude Code, Cursor, Copilot, or equivalent tools as a core part of your daily engineering workflow, have opinions about how they make a team faster, and know how to apply them responsibly when PHI and other sensitive data are in scope.
WhatYou Will Do
Set technical direction across ML training, serving, and observability, and be the final escalation point for the most elusive infrastructure problems (GPU capacity, Spark tuning, production incidents)
Own and evolve our paved-road framework (the shared CI/CD spine, model-workflow scaffolding, and Databricks Asset Bundles) so Data Science teams can go from a config file to a production workflow without bespoke plumbing
Lead architecture for LLM-endpoint serving across managed providers (Databricks, AWS, Snowflake) and self-hosted deployments, covering latency, cost, caching, evaluation, and PHI-safe routing
Own the standards and tooling for MLflow, model registry, training image supply chain, and observability across training and inference
Partner closely with your Data & ML Platform teammates to present a cohesive ML platform to the Data Science, App Dev, and Operations teams at Datavant
Serve as a key technical input to vendor and platform selection decisions across model providers, ML tooling, and observability
Mentor senior engineers on the team, provide technical guidance to platform consumers, and stay hands-on writing high-leverage code and Infrastructure-as-Code alongside your teammates
10+ years of software engineering experience, with 3+ years designing, evolving, and operating enterprise-scale ML platforms in production
Strong technical judgment under ambiguity and a track record of setting standards, influencing peers, and raising the bar across teams
Hands-on production experience with Databricks and/or Amazon Sage Maker, MLflow (or an equivalent tracking + registry system), and at least one core ML framework (PyTorch, Tensor Flow, or similar)
Fluency in Java (or a JVM equivalent) and Python, with real depth in Apache Spark for large-scale data and distributed compute
Real depth in AWS: networking, IAM, GPU compute, and the storage and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).