Founding Computational Scientist, Proteomics Onsite (Cambridge, MA)
Listed on 2026-09-12
-
Research/Development
Data Scientist, Research Scientist
About the Deliverome Project
Deliverome Bio is a nonprofit startup building a first-of-its-kind open atlas of surface protein abundance and internalization. We9re hiring a founding team to generate the data that could expand what targeted therapies can reach.
The RoleYou will build and own the computational side of our proteomics platform: the pipelines, analysis, and data infrastructure that turn raw spectra from surface-enrichment experiments into quantitative measurements of the human surfaceome that we can compare across experiments and release publicly. This role is primarily scientist-facing. The people at the bench are who you serve first, and much of the work is building QC, reports, and result formats they can use directly without waiting on you.
You will also contribute to our broader computational and data infrastructure, including how samples, metadata, and files are tracked, because the person who understands the data model should help decide how the data gets recorded in the first place.
We have an Evosep LC and an Orbitrap Astral Zoom, and we plan to run them at high sample throughput across many cell types and conditions. The data volumes and reprocessing demands that come with that need to be planned for from the beginning rather than handled once we are behind. We are flexible on whether compute lives in the cloud, on local hardware, or both.
What we care about is scale, reproducibility a year later, and being able to release everything we generate. We currently run Spectronaut and are open to changing or adding to it, and we would value your perspective on the tradeoffs between commercial and open-source stacks at this throughput.
We also need someone who can diagnose the full workflow, not only the analysis. Surface proteomics can fail at labeling chemistry, enrichment specificity, digestion, chromatography, acquisition, search parameters, normalization, and statistics. We are looking for someone who can review a set of runs, say where the problem most likely sits, and propose the experiment that distinguishes between the remaining possibilities. That means knowing the wet lab well enough to push on experimental design while it is being decided, not after the samples are made.
This is a rare opportunity to build a public data resource from the first sample onward. No standard exists for how surfaceome abundance and internalization data should be quantified, QC d, and compared across cell types and labs, and you will have unusual latitude to set one. Everything you build, including pipelines, spectral libraries, QC metrics, benchmarks, and the atlas itself, will be openly released and used by drug developers and academic labs worldwide.
Like all of our founding hires, you will have a direct line to the co-founders and a genuine say in how the platform is built. You will gain practical, shareable expertise across large-scale MS analysis, research data infrastructure, and open data release that flexibly advances your career in either industry or academia.
Analysis infrastructure at scale
Build, containerize, and maintain pipelines for bottom-up LC-MS/MS analysis, DIA and DDA, that run reproducibly across hundreds to thousands of files
Own the stack decisions: search and quantification software, workflow orchestration, compute and storage architecture, and cost
Build the data model and provenance layer, covering sample and run metadata, parameter capture, versioned reprocessing, and the ability to answer which pipeline version produced a given number a year later
Make reprocessing routine, so that improved methods are applied across the full corpus and not only to new data
Quantification, statistics, and the atlas
Design the quantification strategy for comparability across cell types and experiments, including normalization, batch structure, missing value handling, and protein inference and rollup
Implement differential and ranking statistics with error control that holds up in review and in reuse by others
Move toward absolute or ratio-anchored abundance estimates where the biology calls for it, and be explicit about the assumptions behind any copies-per-cell number
Explore data completeness across cell and tissue types to understand where more data is needed, or when aspects of the atlas are sufficiently statistically empowered
Build the atlas as a versioned, queryable resource with annotation and topology integration, confidence tiers, and a clear separation between what we measured and what we…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).