×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Data Expert, Data Engineer; Clinical & Multimodal Data Integration

Job in South San Francisco, San Mateo County, California, 94083, USA
Listing for: Genentech
Full Time position
Listed on 2026-07-18
Job specializations:
  • IT/Tech
    Data Engineering, Data Scientist, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 119800 - 222400 USD Yearly USD 119800.00 222400.00 YEAR
Job Description & How to Apply Below
Position: Sr. Data Expert, Data Engineer (Clinical & Multimodal Data Integration)

The Position

A healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That’s what makes us Roche.

Advances in AI, data and computational sciences are transforming drug discovery and development. Roche’s Research and Early Development organizations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D, leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The Computational Sciences Center of Excellence (CS CoE) is a strategic, unified group whose goal is to harness the transformative power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and life‑changing medicines for patients worldwide.

The Computational Sciences Center of Excellence (CS CoE) brings together data, AI, and computational expertise to accelerate innovation across gRED and pRED. Within CS CoE, the Data and Digital Catalyst (DDC) organization leads the modernization of our data ecosystem, enabling scalable, data‑driven science.

The Data Capability organization within DDC is responsible for establishing foundational data capabilities, including data connectivity, data compliance, scientific content management and data ingestion, curation, integration, and delivery. The team ensures that high‑quality, well‑structured datasets are available to power analytics, AI/ML, and scientific discovery across Research and Early Development.

The Opportunity

We are seeking a Sr. Data Expert, Data Engineer to lead the integration and delivery of clinically anchored, multimodal scientific datasets spanning clinical, sequencing, imaging, proteomics, and other emerging data modalities.

  • Lead the integration and harmonization of clinical and multimodal scientific datasets, applying industry data standards and metadata frameworks to improve interoperability and scientific usability.
  • Own end‑to‑end data delivery by designing, validating, documenting, and delivering high‑quality, analysis‑ready datasets that support research, AI/ML, and computational biology initiatives.
  • Develop scalable data workflows that automate data ingestion, quality control, transformation, and metadata management across diverse scientific data sources.
  • Partner with computational scientists, bioinformaticians, and data engineers to understand scientific requirements and translate them into scalable, reusable data solutions.
  • Drive data quality and continuous improvement by implementing validation frameworks, metadata standards, and AI‑assisted data curation practices that improve data discoverability and reuse.
  • Support emerging AI and foundation model initiatives by preparing interoperable, metadata‑rich datasets optimized for downstream analytics and machine‑learning applications.
Who You Are
  • You have a PhD with 2+ years, a Master’s degree with 3–5 years, or a Bachelor’s degree with 5+ years of experience in Bioinformatics, Data Science, Biomedical Engineering, Computer Science, Clinical Sciences, or a related discipline, with experience working with clinical, biomedical, or scientific datasets.
  • You have hands‑on experience integrating clinical data with one or more scientific modalities, including sequencing, imaging, proteomics, or other omics datasets, and understand clinical data models and longitudinal patient data.
  • You are proficient in Python (Pandas), SQL, and scientific data processing, with experience working with scientific data formats such as FASTQ, BAM/CRAM, VCF, DICOM, Ann Data, or Parquet, and familiarity with cloud data platforms (AWS or GCP).
  • You have experience developing or supporting automated data pipelines using workflow orchestration tools such as Airflow, Nextflow, Snakemake, or Prefect, and are comfortable using Git for collaborative software development.
  • You are a collaborative problem solver with a strong focus on data quality, metadata management, and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary