Senior Data Engineer
Listed on 2026-08-13
-
IT/Tech
Data Engineering, Data Analyst
Location:
Hybrid, Iselin, NJ
Medidata follows a hybrid office policy in which employees who are hired for an in-person position are expected to work on site a certain number of days per week following Company policy.
About our Company:Medidata is powering smarter treatments and healthier people through digital solutions to support clinical trials. Celebrating over 25 years of ground-breaking technological innovation across more than 38,000 trials and 12 million patients, Medidata offers industry-leading expertise, analytics-powered insights, and one of the largest clinical trial data sets in the industry. More than 1 million registered users across approximately 2,300 customers trust Medidata's seamless, end-to-end platform to improve patient experiences, accelerate clinical breakthroughs, and bring therapies to market faster.
A Dassault Systemes brand (Euronext Paris: FR0014003TT8, DSY.PA), Medidata is headquartered in New York City and has been recognized as a Leader by Everest Group and IDC. Discover more at Listen to our latest podcast, from Dreamers to Disruptors, and follow us at @Medidata.
Our Data Ops team empowers data-driven features across Factorial by providing advanced ingestion services (SaaS APIs, DBs, Streams), storage capabilities (Data Lake, Stream Storage), data quality, and governance.
As a Senior Data Engineer, you will expand this mission by reinforcing our analytics realm. You will leverage data engineering best practices to implement and maintain robust, scalable data pipelines and processes. You will be part of a small, high-agency team driving the design and operation of data infrastructure that supports product developers, making data a core asset that powers both internal analytics and user-facing product features.
We predominantly use Python and operate with a Git Ops-driven mindset, running short development cycles with ephemer ial dev environments to test, fail fast, and iterate quickly.
Responsibilities:- Take end-to-end ownership of Lakehouse components from design to deployment, applying best practices in batch and streaming, and partnering directly with Infrastructure to operate scalable analytical systems.
- Collaborate with Data Ops and Product teams to integrate data from various sources, design efficient data flows, and enhance our existing analytics platform with robust dashboards and self-service capabilities.
- Write RFCs, lead technical conversations, and share knowledge while mentoring less-experienced engineers, especially Analytics Engineers working within the Lakehouse framework.
- Engage with stakeholders from managers to product teams, translating monitoring and analytics requirements into actionable projects with high-quality deliverables.
- Active participation in team rituals, pairing with different domains to understand their data challenges, and continuously experimenting to improve how data is ingested, curated, and exposed.
- Bachelor's or Master's degree in Computer Science, Information Technology, or a related field with 5+ years of experience.
- Solid experience building and operating reliable data pipelines, with a strong focus on analytics transformation and Lakehouse best practices.
- Experience working with distributed processing and query engines such as Apache Spark and Trino.
- Proven track record of implementing dbt in production environments, including advanced modeling, testing, and documentation.
- Hands-on experience working with Apache Iceberg in production environments.
- Experience with Iceberg catalogs such as Polaris, Lake keeper, or similar solutions for table governance and metadata management.
- Experience with data ingestion tools such as Fivetran, Airbyte, or similar ELT/ETL platforms.
- Familiarity with AWS cloud services.
- Comfort working entirely in English and collaborating within distributed teams.
- Strong knowledge of Terraform and infrastructure concepts, including experience with IaC and cloud resource provisioning.
- Experience designing and implementing Generative AI agents, including hands-on experience with frameworks such as Lang Chain.
As with all roles, Medidata sets ranges based on a number of…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).