Machine Learning & Data Operations Engineer
Listed on 2026-08-31
-
IT/Tech
Data Engineering, Data Scientist, Machine Learning/ ML Engineer
Machine Learning & Data Operations Engineer
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve.
This is hard, urgent, selfless work—but it's work worth doing. If you're driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
As a Machine Learning & Data Operations Engineer on Tune Lab, you will build cutting-edge ML and AI tools alongside a team of engineers and scientists to accelerate and enhance Lilly's drug discovery process. You will take a hands-on role across the full lifecycle of models and the data that feeds them: moving trained models from research into reliable production environments, running inference at scale, and building the data pipelines and readiness checks that keep the data substrate underpinning those models trustworthy.
You will stand up the validation, monitoring, and model-card review that keep both models and data production-ready—catching anomalies, schema drift, and performance regressions before they reach researchers. You will collaborate closely with partners across Lilly Research Labs, AI, Software Engineering, Data Science, and IT Operations, along with industry-leading external collaborators, to put the power of ML and computational tooling directly into researchers' day-to-day work.
Responsibilities
Model Deployment, Serving & Inference
- Move trained models from research and experimentation into production, packaging, versioning, and promoting them across development, staging, and production environments and across cloud targets (AWS, Azure, GCP) and on-prem or hybrid infrastructure
- Build and operate scalable inference services and APIs—batch, real-time, and streaming—delivering low-latency, high-throughput serving that meets researcher and downstream-system needs
- Design and maintain model-serving infrastructure using containers and Kubernetes, with autoscaling, versioned rollouts (e.g., blue-green or canary), and rollback so updates ship without disrupting users
- Integrate models into researcher-facing tools and enterprise systems, ensuring seamless interoperability and data flow across platforms
Data Pipelines & Readiness
- Design, build, and maintain scalable, secure data pipelines—batch, change-data-capture (CDC), and streaming—that move and transform data across the platform, including the embedding, vectorization, and feature pipelines that feed downstream ML and LLM applications
- Implement scalable storage and retrieval for large-scale structured and unstructured scientific data across cloud and on-prem or hybrid infrastructure
- Build and operate automated data-readiness and quality-monitoring workflows for high-dimensional scientific and enterprise datasets, including multi-method anomaly and outlier detection across numerical and categorical data
- Validate files for missing values, illegal characters, and structural issues, and build schema-drift detection with historical tracking and automated reporting—catching data-contract changes before they reach models and significantly reducing manual data QA
Model & Data Validation, Monitoring & Governance
- Author, review, and validate model cards—verifying documented performance, intended use, limitations, data lineage, and evaluation results before models are promoted
- Run and automate model validation and evaluation—reproducing metrics, checking calibration and performance against acceptance criteria, and gating promotion on the results
- Implement production monitoring for model, data, and service health—latency, throughput, data and prediction drift, and quality—with alerting and proactive remediation
- Define acceptance criteria, audit trails, and reproducible checks; adjudicate flagged data and model issues with data owners and scientists; and track and report…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).