Machine Learning Engineer, Infra, AI for Drug Discovery
Listed on 2026-08-03
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
The Position
A healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That’s what makes us Roche.
Advances in AI, data, and computational sciences are transforming drug discovery and development. Roche’s Research and Early Development organisations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D, leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The new Computational Sciences Center of Excellence (CoE) is a strategic, unified group whose goal is to harness this transformative power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and transformative medicines for patients worldwide.
TheOpportunity
At Roche’s AI for Drug Discovery (AI4DD) group (Prescient Design), we are building the machine learning platforms that enable researchers and engineers to move models from experimentation into reliable scientific and production workflows. We are seeking a Machine Learning Infrastructure Engineer to help build and operate the platforms that support model deployment, evaluation, promotion, monitoring, and lifecycle management across the organization.
This role will contribute to our model-serving platform, and to the broader infrastructure required to make machine learning models easier to deploy, scale, observe, and safely incorporate into scientific and agentic workflows.
The scope extends beyond LLM serving. You will work with a range of machine learning and scientific models, including real-time and batch inference workloads, GPU-backed services, agentic applications, and our in-silico drug discovery workflows. This is a hands‑on engineering role for someone who enjoys writing and shipping production software across application code, cloud infrastructure, Kubernetes, and distributed systems. Prior inference‑platform experience is helpful but not required;
prior experience in biotech or drug discovery is also helpful but not required; we value strong engineering fundamentals, curiosity, and the ability to take platform problems from design through production operation.
- Design, implement, ship, and operate scalable model‑serving infrastructure for machine learning, scientific, LLM, and agentic workloads.
- Help evolve our internal model deployment platform into a reliable, self‑service platform for teams across the organization.
- Improve platform scalability and reliability, including scale‑to‑zero, faster model startup, workload isolation, traffic management, and reduction of request failures and latency bottlenecks.
- Build observability and operational tooling for model usage, latency, reliability, resource consumption, inference cost, bottlenecks, and service‑level indicators.
- Improve the usability of model deployment by developing validated configuration interfaces, reusable deployment patterns, APIs, command‑line tools, and documentation.
- Help converge real‑time and batch inference workflows onto shared platform capabilities where appropriate.
- Contribute to model lifecycle management infrastructure, including model registration and versioning, evaluation, promotion and release gates, monitoring, environment progression, and rollback.
- Build event‑driven integrations that connect model publication, evaluation, promotion, deployment, and retraining workflows.
- Build consistent metrics and evaluation signals for understanding model cost, quality, reliability, and fitness for downstream workflows.
- Partner with machine learning, data, scientific, and platform teams to translate requirements into maintainable solutions and remove infrastructure bottlenecks.
- Own work streams from design through implementation and production support, using strong software‑engineering practices including testing, reviews, documentation, and incremental delivery.
- BS or MS in Computer…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).