More jobs:
Senior Scientific Data Curator
Job in
Indianapolis, Marion County, Indiana, 46201, USA
Listed on 2026-08-31
Listing for:
Eli Lilly
Full Time
position Listed on 2026-08-31
Job specializations:
-
IT/Tech
Data Engineering, Data Scientist
Job Description & How to Apply Below
Senior Scientific Data Curator
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve.
This is hard, urgent, selfless work—but it's work worth doing. If you're driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
- Conduct comprehensive inventory of historical and ongoing datasets across Tune Lab's modeling domains—small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo PK and toxicology, antibody and biologics develop ability, and clinical PK/PD—spanning therapeutic areas (oncology, immunology, metabolic diseases, neuroscience, etc.) and 20+ years of discovery, preclinical, and clinical data
- Assess and score data quality, completeness, and integration feasibility for each dataset, accounting for the distinct data structures of each domain, including assay and dose–response measurements, concentration–time profiles and dosing regimens, in vivo study readouts, sequence- and structure-derived features for biologics, and biomarker and clinical covariate data
- Map metadata gaps across legacy systems and source platforms, documenting study contexts, assay and protocol methods, protocol deviations, data quality flags, and provenance information
- Develop automated pipelines (including LLM-assisted extraction where appropriate) to identify and extract domain-relevant data from internal documents, assay databases, and study reports into standardized, model-ready formats
- Produce a prioritized data assessment report recommending which domains, therapeutic areas, and indications to integrate first, based on data volume, complexity, portfolio relevance, and model feasibility
- Design and implement standardized, extensible schemas for the integrated multi-domain database, working with computational partners to ensure AI/ML readiness across small-molecule and biologics modalities
- Build and maintain data harmonization pipelines: label normalization, unit and assay-condition standardization across studies and sources, time-point and dose alignment, sequence and structure normalization for biologics, and covariate encoding
- Apply domain-driven quality control practices—sequence validation, hidden duplicate detection, cross-source discrepancy resolution, and cross-species dataset integration using allometric scaling where applicable
- Develop and execute outlier detection protocols, flagging and adjudicating anomalous values in collaboration with clinical pharmacologists, DMPK scientists, toxicologists, and antibody engineers as appropriate to the domain
- Create reproducible data quality assurance workflows with documented acceptance criteria and audit trails
- Curate and enrich metadata to enable cross-study and cross-domain querying—linking compound and molecule identifiers, sequence and construct identifiers, assay methods, formulation details, and study design parameters
- Serve as the primary data domain expert for external consortium partners working within Lilly's controlled cloud environment
- Collaborate with pharmacometricians, DMPK scientists, toxicologists, and antibody engineers to validate harmonized datasets against legacy models and established analyses (e.g., NONMEM/Monolix outputs for the clinical PK/PD domain)
- Work with the Tune Lab ML team to ensure curated datasets meet the input specifications for the platform's multi-task ML models, representation and foundation-model embeddings, and mechanistic/hybrid PK/PD frameworks (e.g., Neural ODE, SINDy)
- Contribute to platform deployment by supporting the development of data dictionaries, user documentation, and training materials for internal and consortium end users
- M.S. or…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×