Data Platform Engineer
Listed on 2026-10-01
-
Software Development
Data Engineering
Final date to receive applications: 19 October 2026
Department: Informatics
Employment Type: Full Time
Location: Boston
Compensation: $115,000 - $130,000 / year
DescriptionWe are hiring a Data Platform Engineer to build and operate the pipelines that take raw data through to analysis-ready datasets, and to own the datasets those pipelines produce. Today that means running orchestration over cloud compute, supporting both CRO deliveries and wet-lab experiment cycles.
You will sit between the wet lab and CROs who generate our data and the computational biology and ML teams who consume it, and you will make training and inference runs traceable enough to explain months after they happened.
We're hiring one person for this role, and they can be based in either London or Boston.
Here is our timeline for hiring this role:
- Now until October 19th: Accepting applications
- October 26th:
Planned start for interviews
- Build and operate the pipelines that take raw data through to analysis-ready datasets. Routine runs should complete without someone having to watch or shepherd them.
- Own the datasets those pipelines produce: schema stability, validation, provenance, versioning and documentation, along with the transformation layer that turns processed outputs into tables people can actually query. Scientists and downstream systems should be able to rely on an output without first checking what changed upstream.
- Work with stakeholders on either side of the data, the wet lab and CROs who generate it and the computational biology and ML teams who consume it. That means understanding how the data is produced and what it is used for. When a problem recurs, sometimes the right fix is in the pipeline and sometimes it is in how the data is produced or delivered.
- Maintain good engineering practice across the dry-lab codebase: useful code review, meaningful tests, CI, and failures that are visible and diagnosable.
- Make training and inference runs traceable through versioned inputs, recorded configuration, and artifacts that can be tied back to the data and code that produced them, so that an important run can be explained months after it happened.
- Three to five years building and operating production data infrastructure, mainly in Python. You have owned pipelines that other people depended on and dealt with them when they failed.
- You have worked in a team with solid engineering practice and know what good review, testing and deployment look like day to day.
- Comfortable with a workflow orchestrator such as Dagster, Airflow or Prefect, and with configuring cloud compute directly. We run on AWS, so Batch, Fargate and S3 experience matters.
- You can work from a specification, and you will call out gaps or bad assumptions rather than quietly implementing them.
- You can talk to a scientist about how their data is generated, understand the practical constraints, and tell those apart from preferences or one-off requests.
- You are motivated by making systems reliable and maintainable.
- Nice to have: biological or scientific data, particularly sequencing or omics, including the awkward file formats and incomplete metadata that come with it;
Dagster in production and infrastructure-as-code with Terraform or equivalent; analytical stores such as Click House or DuckDB and transformation tooling such as dbt or SQLMesh; working closely with a wet lab, or with data whose quality depends partly on what happens at the bench; building data infrastructure for LLM or agentic systems, where changes to schemas, metadata or provenance can silently affect the output.
Join Outpost Bio?
- You'll own real equity in what you build. We offer meaningful stock options because we believe the people building this company should share in what it becomes. We want teammates who think like owners, and we structure compensation to reflect that.
- Outstanding benefits. Full medical, dental and vision from day one, with Outpost covering 100% of the employee premium on the base plan and 50% for dependents. 401(k) with a 3% match. Short and long-term disability, employer paid. 50% of your MBTA Perq commuter pass. 25 days PTO plus your birthday off, and a paid winter break between Christmas Eve and New Year.
- An ML Lab-in-the-Loop. Your work feeds directly into Outpost's AI platform, and the platform feeds back into the next experiment. The loop between the wet lab, the data and the models runs in days, not years, and you'll iterate inside it whichever side you…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).