Senior Data Scientist, Biologics Discovery
Listed on 2026-09-03
-
IT/Tech
Data Scientist, Machine Learning/ ML Engineer, Data Engineering, AI Engineer (Applied/Software)
At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and Med Tech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity.
Learn more at
As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit.
Job FunctionData Analytics & Computational Sciences
Job Sub FunctionData Science
Job CategoryScientific/Technology
All Job Posting LocationsMadrid, Spain, Raritan, New Jersey, United States of America, Spring House, Pennsylvania, United States of America, Titusville, New Jersey, United States of America
Job DescriptionOur expertise in Innovative Medicine is informed and inspired by patients, whose insights fuel our science-based advancements. Visionaries like you work on teams that save lives by developing the medicines of tomorrow.
Join us in developing treatments, finding cures, and pioneering the path from lab to life while championing patients every step of the way.
Learn more at
About the opportunityJohnson & Johnson Innovative Medicine is seeking a Senior Data Scientist dedicated to our Biologics Discovery organization. This role sits within our Data, Data Science & Artificial Intelligence team (DDSAI) and partners closely with our In Silico Discovery (ISD) organization - the group that builds the molecular design and property-prediction models (for example, develop ability, affinity and binding, and other molecular-property and liability-risk models) that guide which biologic molecules to design, make, and advance.
ISD owns core molecular model development; you will build the data-facing ML capabilities (featurization, model-ready datasets, evaluation frameworks, and applied models on assay and sequence data) that make ISD's models faster to build and better to trust.
This position will be based at one of our office locations in either Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ, USA; or Madrid, Spain. (No remote option.)
Please note that this role is available across multiple countries and may be posted under different requisition numbers to comply with local requirements. While you are welcome to apply to any or all of the postings, we recommend focusing on the specific country(s) that align with your preferred location(s).
USA
- Requisition Number: R-095854
Spain
- Requisition Number: R-096793
Why this role matters:
Biologics Discovery is generating rich, fast-growing data across assays, sequences, and modalities, and the opportunity now is to make that data fully model-ready and seamlessly available for ML. This role ensures biologics data is structured for training, and that applied ML on discovery data helps scientists prioritize molecules, flag risks, and generate hypotheses earlier - strengthening the interface to ISD's models rather than duplicating them.
Position Summary
You will design robust featurization and dataset curation, build and evaluate applied models on biologics assay, biophysical, and sequence/construct data, and define evaluation frameworks that keep models trustworthy. You operate at the interface between our data-generating and data-infrastructure partners and In Silico Discovery (ISD), ensuring the datasets and features you create strengthen ISD's molecular property models. This is an opportunity to shape how AI learns from every biologics experiment.
Key Responsibilities:Featurization & Model-Ready Data
- Develop featurization and model-ready datasets from antibody/protein sequence, construct, assay, and biophysical data.
- Work with data engineers to specify the features, labels, and levels of aggregation that models need, preserving raw representations where information matters.
- Curate, document, and version datasets so modeling is reproducible and traceable.
- Develop featurization and model-ready datasets from antibody/protein sequence, construct, assay, and biophysical data.
- Work with data engineers to specify the features, labels, and levels of aggregation that models need, preserving raw representations where information matters.
- Curate, document, and version datasets so modeling is reproducible and traceable.
- Collaborate with ISD to hand off standardized, traceable training datasets and align on where Data Science enables versus where ISD owns modeling.
- Partner with Discovery scientists to frame ML problems around real decision points in the design-make-test-learn (DMTL) cycle.
- Work closely with ontology and MLOps…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).