×
Register Here to Apply for Jobs or Post Jobs. X

Senior Data Scientist, Biologics Discovery

Job in Raritan, Somerset County, New Jersey, 08869, USA
Listing for: J&J Family of Companies
Full Time position
Listed on 2026-09-03
Job specializations:
  • IT/Tech
    Data Scientist, Machine Learning/ ML Engineer, Data Engineering, AI Engineer (Applied/Software)
Job Description & How to Apply Below

Senior Data Scientist

Our expertise in Innovative Medicine is informed and inspired by patients, whose insights fuel our science-based advancements. Visionaries like you work on teams that save lives by developing the medicines of tomorrow.

Join us in developing treatments, finding cures, and pioneering the path from lab to life while championing patients every step of the way.

Johnson & Johnson Innovative Medicine is seeking a Senior Data Scientist dedicated to our Biologics Discovery organization. This role sits within our Data, Data Science & Artificial Intelligence team (DDSAI) and partners closely with our In Silico Discovery (ISD) organization - the group that builds the molecular design and property-prediction models that guide which biologic molecules to design, make, and advance.

ISD owns core molecular model development; you will build the data-facing ML capabilities that make ISD's models faster to build and better to trust.

This position will be based at one of our office locations in either Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ, USA; or Madrid, Spain. (No remote option.)

Why this role matters:
Biologics Discovery is generating rich, fast-growing data across assays, sequences, and modalities, and the opportunity now is to make that data fully model-ready and seamlessly available for ML. This role ensures biologics data is structured for training, and that applied ML on discovery data helps scientists prioritize molecules, flag risks, and generate hypotheses earlier - strengthening the interface to ISD's models rather than duplicating them.

Position Summary

You will design robust featurization and dataset curation, build and evaluate applied models on biologics assay, biophysical, and sequence/construct data, and define evaluation frameworks that keep models trustworthy. You operate at the interface between our data-generating and data-infrastructure partners and In Silico Discovery (ISD), ensuring the datasets and features you create strengthen ISD's molecular property models. This is an opportunity to shape how AI learns from every biologics experiment.

Key Responsibilities:

  • Develop featurization and model-ready datasets from antibody/protein sequence, construct, assay, and biophysical data.
  • Work with data engineers to specify the features, labels, and levels of aggregation that models need, preserving raw representations where information matters.
  • Curate, document, and version datasets so modeling is reproducible and traceable.
  • Develop featurization and model-ready datasets from antibody/protein sequence, construct, assay, and biophysical data.
  • Work with data engineers to specify the features, labels, and levels of aggregation that models need, preserving raw representations where information matters.
  • Curate, document, and version datasets so modeling is reproducible and traceable.
  • Collaborate with ISD to hand off standardized, traceable training datasets and align on where Data Science enables versus where ISD owns modeling.
  • Partner with Discovery scientists to frame ML problems around real decision points in the design-make-test-learn (DMTL) cycle.
  • Work closely with ontology and MLOps colleagues so datasets carry consistent semantics and models move reliably from development into use.
  • Champion reproducibility, documentation, and responsible AI.

Why This Role Is Unique

This is an opportunity to apply ML where it truly moves the needle in biologics discovery - grounded in real assay and sequence data, tightly partnered with world-class molecular modeling, and with real room to grow your scope, technical leadership, and impact as you build a track record of delivery.

Qualifications:

Required

  • Master's or Ph.D. in Computer Science, Machine Learning, Computational Biology, Bioinformatics, Statistics, or a related field.
  • At least 2 years of applied ML experience, including model development, evaluation, and dataset curation on complex scientific or biomedical data.
  • Strong proficiency with Python and the modern ML stack (e.g., PyTorch, scikit-learn) and SQL.
  • Experience turning complex, heterogeneous experimental data into robust features and training sets, with exposure to cloud training and data infrastructure.
  • Sound understanding of evaluation, validation, and the risks of leakage and distribution shift.
  • Ability to collaborate effectively with experimental scientists and modeling partners in a matrixed R&D environment.

Preferred

  • Experience with biologics, antibody/protein sequence models, or protein language models.
  • Experience with active learning, Bayesian optimization, or sequence-based generative models for molecular design.
  • Familiarity with biophysical/assay data and develop ability endpoints.
  • Experience with MLOps, experiment tracking, and model monitoring.
  • Familiarity with how ontologies or knowledge graphs support data reuse and AI-ready datasets.
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary