Senior Data Engineer — Translational Data Products
Listed on 2026-08-29
-
IT/Tech
Data Engineering
Working with Us Challenging. Meaningful. Life-changing. Those aren’t words that are usually associated with a job. But working at Bristol Myers Squibb is anything but usual. Here, uniquely interesting work happens every day, in every department. From optimizing a production line to the latest breakthroughs in cell therapy, this is work that transforms the lives of patients, and the careers of those who do it.
You’ll get the chance to grow and thrive through opportunities uncommon in scale and scope, alongside high-achieving teams. Take your career farther than you thought possible.
Bristol Myers Squibb recognizes the importance of balance and flexibility in our work environment. We offer a wide variety of competitive benefits, services and programs that provide our employees with the resources to pursue their goals, both at work and in their personal lives. Read more About The Role Our team builds and maintains data products that make R&D data usable and AI-ready biomarker, biospecimen, clinical trial, omics, among others, assembled into governed, curated, linked, documented, AI-ready products that translational scientists and analysts use for insight generation and decision making.
You will own data products end to end — from documenting where the data lives, to unlocking the source system with its data owner, through entity mapping, modeling and validation, to making sure the product answers high-value scientific questions and gets used or retired. You will work alongside data scientists, machine learning engineers, and translational medicine stakeholders, and you will use Claude Code as a normal part of how you deliver.
What you will do
- Build, operate and own data products. Design, build, test, and maintain ETL/data normalization pipelines on Databricks and AWS that create governed, well-documented datasets in Unity Catalog (both tables and views).
- Unlock data access. Work directly with data owners and governance partners to bring new sources under governance and make them AI-ready — cataloged, described, permissioned, quality- checked, and discoverable. This may include csv, tsv, json, xlsx, power point, structured documents, etc.
- Pursue business impact. Engage stakeholders in translational medicine, biomarker sciences, predictive medicine, reverse translation and clinical development to understand and translate the high value questions they are trying to answer into robust data product designs and data maps linking sources to targets (STTM).
- Translate between business and data. Turn a scientific or business question into a data model and a mapping from the question to the SQL query; carefully describe data constraints / context and considerations into caveats that a stakeholder can understand.
- Validate everything. Bring a test-driven, validation-driven mindset expectations and data quality checks in the pipeline, reconciliation against source, unit and regression tests that run automatically in Git Hub Actions on every pull request, and documented evidence that a number is correct before anyone depends on it.
- Engineer for robustness. Create the hooks, guardrails, CDK / Cloud Formation infrastructure as code, and automated checks that keep our work reliable and aligned with BMS enterprise IT standards — SDLC, security, SSO / identity and access management, mandatory resource tagging and service registration (CRID, ISR assessment), code review, secrets handling, change management. Run the required cyber checks — dependency, container and infrastructure-as-code scanning (Wiz, Dependabot) on every pull request — and respond quickly when a vulnerability is reported, remediating critical findings within BMS Cybersecurity timelines.
- Work with ML engineers. Prepare feature-ready and model-ready datasets, and support ML and GenAI workloads built on top of our products; support deployment and accessibility of data products on NVIDIA clusters.
- Work as part of a team. Git Hub-based development, pull requests, code review, automated build/test/deploy pipelines (continuous integration and delivery, CI/CD), infrastructure as code.
- Continuous improvement. Work with stakeholders to ensure the data products…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).