Principal Scientist, Data Science; Data Products, Integration & Analysis
Listed on 2026-07-20
-
IT/Tech
Data Engineering, Data Scientist, AI Engineer (Applied/Software), Data Analyst
Location: Spring House
Semantic-ready data products
Feature stores
Predictive model development
Position SummaryThe Principal Scientific Data Scientist will lead the design, implementation, and evolution of scientific data products and integration strategies supporting AI‑enabled drug discovery and development.
This individual will create scalable, interoperable, AI‑ready data products that connect discovery, preclinical, clinical, safety, and real‑world evidence domains, enabling the creation of validated‑biomarker data assets. The role will establish the data architecture, integration strategy, metadata framework, and productization approach needed to support semantic reasoning, knowledge graphs, GraphRAG, advanced analytics, and agentic AI applications.
Working closely with scientific stakeholders, knowledge architects, AI engineers, and Amazon Bio Discovery platform teams, this individual will define the future‑state scientific data ecosystem and ensure high‑quality data products are delivered to support translational science and patient safety initiatives.
Build AI reasoning models to support data‑driven translational safety decision making.
MissionBuild and operationalize AI‑ready scientific data products that enable seamless integration, harmonization, and reuse of data across the drug discovery and development lifecycle.
Key Responsibilities- Define and execute a scientific data product strategy supporting discovery research, translational science, preclinical safety, clinical development, pharmacovigilance, and real‑world evidence.
- Establish reusable, scalable data products that support analytics, AI, knowledge graph, and scientific reasoning use cases.
- Develop product roadmaps aligned with organizational priorities and scientific objectives.
- Design integration frameworks connecting heterogeneous scientific data sources.
- Define data harmonization strategies spanning SEND, SDTM, ADaM, MedDRA, Imaging, Omics, Biomarker, Pathology, and real‑world data.
- Create architecture patterns supporting cross‑domain data interoperability.
- Define the implementation strategy for scientific data products deployed on AWS.
- Partner with Amazon engineering and platform resources to deliver scalable data pipelines and products.
- Provide technical leadership and architectural oversight for implementation activities aligned with the Data Strategy group.
- Ensure digital solutions align with enterprise architecture, security, governance, and AI‑readiness requirements.
- Lead design and implementation of curated datasets, semantic‑ready data products, feature stores, metadata products, scientific data services, and AI‑ready data assets.
- Establish reusable patterns for data onboarding, transformation, validation, and publication.
- Define metadata standards and data quality frameworks.
- Implement lineage, provenance, traceability, and FAIR data principles.
- Establish monitoring and quality controls for scientific data products.
- Build predictive AI/ML models to support translational safety decision making.
- Partner with discovery scientists, toxicologists, clinical scientists, safety scientists, data scientists, data strategy business partners, knowledge architects, and AI engineers.
- Translate scientific questions into scalable data products and technical solutions.
- Master’s or PhD in Computer Science, Data Engineering, Bioinformatics, Biomedical Informatics, Information Systems, Computational Biology, or a related scientific discipline.
- 5+ years of experience in scientific data engineering, data architecture, data products, or life sciences informatics.
- Demonstrated experience designing and delivering enterprise‑scale scientific data products.
- Experience supporting drug discovery, development, clinical research, or pharmacovigilance organizations.
- Experience developing predictive models in drug discovery, development, clinical research, or pharmacovigilance organizations.
- Data architecture
- Data modeling
- Data product design
- Cloud‑native data platforms
- Metadata management
- Data governance
- Predictive model development
- AWS‑based data platforms
- Data lakes and lake houses
- Distributed data…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).