×
Register Here to Apply for Jobs or Post Jobs. X

Senior Scientific Data Curator

Job in Indianapolis, Hamilton County, Indiana, 46262, USA
Listing for: Initial Therapeutics, Inc.
Full Time position
Listed on 2026-09-03
Job specializations:
  • IT/Tech
    Data Scientist, Data Analyst
  • Research/Development
    Data Scientist
Salary/Wage Range or Industry Benchmark: 132000 - 244000 USD Yearly USD 132000.00 244000.00 YEAR
Job Description & How to Apply Below
Location: Indianapolis

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve.

This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Position Summary

The Senior Scientific Data Curator will lead the systematic discovery, assessment, harmonization, and quality assurance of Lilly’s scientific datasets across the full breadth of Tune Lab’s modeling domains—small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo pharmacokinetics and toxicology, antibody and biologics develop ability, and clinical PK/PD—in support of a strategic, cross-modality data unification initiative. This role sits at the intersection of biological and pharmacological domain expertise and data science, translating decades of fragmented, heterogeneous datasets spanning discovery through the clinic into a unified, AI-ready data infrastructure.

The curator will partner closely with computational scientists, DMPK scientists, pharmacometricians, antibody engineers, and external consortium collaborators to ensure that the data substrate underpinning Tune Lab’s federated AI/ML models is comprehensive, well-documented, and scientifically sound.

Core Responsibilities
  • Conduct comprehensive inventory of historical and ongoing datasets across Tune Lab’s modeling domains—small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo PK and toxicology, antibody and biologics develop ability, and clinical PK/PD—spanning therapeutic areas (oncology, immunology, metabolic diseases, neuroscience, etc.) and 20+ years of discovery, preclinical, and clinical data
  • Assess and score data quality, completeness, and integration feasibility for each dataset, accounting for the distinct data structures of each domain, including assay and dose–response measurements, concentration–time profiles and dosing regimens, in vivo study readouts, sequence- and structure-derived features for biologics, and biomarker and clinical covariate data
  • Map metadata gaps across legacy systems and source platforms, documenting study contexts, assay and protocol methods, protocol deviations, data quality flags, and provenance information
  • Develop automated pipelines (including LLM-assisted extraction where appropriate) to identify and extract domain-relevant data from internal documents, assay databases, and study reports into standardized, model-ready formats
  • Produce a prioritized data assessment report recommending which domains, therapeutic areas, and indications to integrate first, based on data volume, complexity, portfolio relevance, and model feasibility
Data Harmonization & Integration
  • Design and implement standardized, extensible schemas for the integrated multi-domain database, working with computational partners to ensure AI/ML readiness across small-molecule and biologics modalities
  • Build and maintain data harmonization pipelines: label normalization, unit and assay-condition standardization across studies and sources, time-point and dose alignment, sequence and structure normalization for biologics, and covariate encoding
  • Apply domain-driven quality control practices—sequence validation, hidden duplicate detection, cross-source discrepancy resolution, and cross-species dataset integration using allometric scaling where applicable
  • Develop and execute outlier detection protocols, flagging and adjudicating anomalous values in collaboration with clinical pharmacologists, DMPK scientists, toxicologists, and antibody engineers as appropriate to the domain
  • Create reproducible data quality assurance workflows with documented acceptance criteria and audit trails
  • Curate and enrich metadata to enable cross-study and cross-domain querying—linking compound and…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary