×
Register Here to Apply for Jobs or Post Jobs. X

Data Engineer II- Life Sciences

Job in New York, New York County, New York, 10261, USA
Listing for: h1
Full Time position
Listed on 2026-09-20
Job specializations:
  • Software Development
    Python, Data Engineering
Salary/Wage Range or Industry Benchmark: 110000 - 135000 USD Yearly USD 110000.00 135000.00 YEAR
Job Description & How to Apply Below
Location: New York

At H1, we believe access to the best healthcare information is a basic human right. Our mission is to provide a platform that can optimally inform every doctor interaction globally. This promotes health equity and builds needed trust in healthcare systems. To accomplish this, our teams harness the power of data and AI-technology to unlock groundbreaking medical insights and convert those insights into action that result in optimal patient outcomes and accelerates an equitable and inclusive drug development lifecycle.

Visit  to learn more about us.

As part of H1’s hiring process, all candidates are required to participate in an in‑person final interview. Depending on your location, this may require travel.

H1's Data Network (H1DN) team is the client‑data mastering network at the core of how H1's products get their data. We run production ingestion for major enterprise customers. Clinical trial data is one of our highest‑visibility streams: it feeds decisions about where trials run and who runs them, and the people who depend on it are as often clinical experts as they are engineers.

SLAs and customer expectations drive how we work, and we're looking for engineers who are energized by that.

WHAT YOU'LL DO AT H1

As a Data Engineer II on the H1DN team, you will build and operate the pipelines behind H1's clinical trials data. You'll work primarily in Python, PySpark, and SQL, and you'll work directly with clinical subject matter experts and Customer Success Managers to turn their domain knowledge into pipeline logic that holds up in production.

  • Build and maintain the Python and PySpark pipelines behind the CTMS trial data pipeline intake workflows, including scoring and status logic.
  • Develop the transformation logic that maps raw trial and customer data to H1's internal data models, handling diverse source formats including CSV, JSON, Parquet, and APIs.
  • Write and tune SQL against large datasets to investigate data questions, validate pipeline output, and support analysis that clinical SMEs and customer‑facing teams depend on.
  • Turn around customer‑driven changes quickly, scoping requests as they arrive, shipping changes that hold up under enterprise SLAs, and reworking logic as customer needs shift mid‑flight.
  • Partner with clinical SMEs to translate domain expertise into concrete data rules, then walk them through the results, explain what the pipeline did and why, and fold their feedback back into the logic.
  • Build the data quality checks, validation logic, and reconciliation that let non‑engineers trust pipeline output without reading the code.
  • Participate in code reviews, maintaining a high bar for quality and adherence to engineering standards.
  • Monitor and improve pipeline observability, contributing to alerting and dashboards that surface job health and data anomalies for both the team and internal users.
ABOUT YOU

You are a data engineer with a strong Python foundation and real distributed‑processing experience. You're drawn to high‑impact teams where the work is tangible: pipelines running, enterprise customers getting their data on time, clinical data that people make real decisions from. You're comfortable in an environment where recurring production runs and customer SLAs shape day‑to‑day priorities, and where a customer request can reorder your week.

You'd rather sit down with a domain expert and understand why the data looks the way it does than build to a spec handed to you secondhand.

You bring experience:

  • Building and shipping production data pipelines in Python, with an understanding of what makes them reliable and maintainable under real load
  • Working with PySpark or a comparable distributed processing framework on datasets too large for a single machine
  • Writing SQL well enough to answer…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary