×
Register Here to Apply for Jobs or Post Jobs. X

Data Engineer, Clinical Operations

Job in Princeton, Mercer County, New Jersey, 08543, USA
Listing for: Bristol Myers Squibb EU Policy
Full Time position
Listed on 2026-07-23
Job specializations:
  • Software Development
    Data Engineering, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 87810 - 106399 USD Yearly USD 87810.00 106399.00 YEAR
Job Description & How to Apply Below

Working with Us
Challenging. Meaningful. Life‑changing. Those aren’t words that are usually associated with a job. But working at Bristol Myers Squibb is anything but usual. Here, uniquely interesting work happens every day, in every department. From optimizing a production line to the latest breakthroughs in cell therapy, this is work that transforms the lives of patients and the careers of those who do it.

You’ll get the chance to grow and thrive through opportunities uncommon in scale and scope, alongside high‑achieving teams. Take your career farther than you thought possible.

Position Summary

As a Data Engineer
, you will support the broader Data Engineering community to deliver cutting‑edge data and analytics platforms for the Global Drug Development (GDD) IT group within the Cross Study Operations and Specimen Management domain. You will design, build, and maintain scalable, production‑grade data pipelines and platform components that support cross‑study data aggregation, specimen tracking, and biobanking workflows. The role requires deep expertise in data engineering, cloud platforms, and Generative AI to solve complex clinical data challenges.

Key Responsibilities
  • Collaborate with BI&T partners, cross‑study operations leads, specimen management specialists, clinical study teams, and cross‑functional leaders to support adoption of the Data Platform.
  • Design, build, and maintain production‑grade data pipelines and platform components for Cross Study Operations and Specimen Management product lines.
  • Develop and enhance data solutions to accelerate data usage across cross‑study clinical R&D programs, ensuring robustness, interoperability, and scalability.
  • Optimize data platform components for performance, scalability, and cost‑effectiveness using techniques such as cloud‑native parallel processing, Databricks Delta Lake, caching, and partitioning.
  • Help design scalable ETL/ELT pipelines and data models using Databricks, Delta Lake, cloud‑native tools, semantic modeling, and interoperability standards for large, complex datasets.
  • Partner with business and data product owners to deliver hands‑on technical solutions for data product development, standardization, testing, lineage, and governance.
  • Utilize Databricks Unity Catalog to enforce data governance, manage metadata, and ensure end‑to‑end data lineage.
  • Contribute to the development of self‑service data discovery solutions to improve findability and reusability of operational and specimen data assets.
  • Maintain thorough documentation of processes, data structures, and technical solutions; provide clear technical recommendations.
  • Help build and deploy GenAI‑powered and NLP‑driven applications that deliver measurable outcomes across operations and specimen management.
  • Develop and operationalize cloud‑based GenAI and LLM‑powered applications using RAG, fine‑tuning, and vector embeddings.
  • Leverage Databricks Mosaic AI and MLflow to develop, track, deploy, and manage machine learning and GenAI models at scale.
  • Stay current on technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization.
  • Serve as a go‑to technical expert, providing guidance and mentorship to junior analysts, interns, and vendor resources.
  • Foster a culture of continuous learning, engineering excellence, and knowledge sharing across the data engineering community.
Qualifications & Experience
  • 5+ years of hands‑on experience in Data Engineering, Analytics, and AI/ML in a cloud environment.
  • Hands‑on expertise with Databricks (Delta Lake, Unity Catalog, Workflows, Mosaic AI, MLflow). Databricks certification is a plus.
  • Strong proficiency with cloud‑native data platforms, ETL/ELT pipeline design, data modeling, and semantic analytics for large datasets.
  • Proficiency in Python, SQL, Spark (PySpark on Databricks), and GenAI frameworks; experience with LLM architectures, RAG, and prompt engineering.
  • Experience developing production‑grade GenAI applications, predictive models, and self‑service analytics tools.
  • Strong stakeholder engagement and communication skills; ability to explain technical concepts to non‑technical audiences.
  • Commitment to engineering excellence,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary