×
Register Here to Apply for Jobs or Post Jobs. X

Principal Data Engineer

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: Dassault Systèmes
Full Time position
Listed on 2026-09-06
Job specializations:
  • IT/Tech
    Data Engineering, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 120000 - 180000 GBP Yearly GBP 120000.00 180000.00 YEAR
Job Description & How to Apply Below
Location: Greater London

Location: This is a hybrid remote/in-office role.

About our Company:

Medidata is powering smarter treatments and healthier people through digital solutions to support clinical trials. Celebrating over 25 years of ground-breaking technological innovation across more than 38,000 trials and 12 million patients, Medidata offers industry-leading expertise, analytics-powered insights, and one of the largest clinical trial data sets in the industry. More than 1 million registered users across approximately 2,300 customers trust Medidata's seamless, end-to-end platform to improve patient experiences, accelerate clinical breakthroughs, and bring therapies to market faster.

A Dassault Systèmes brand (Euronext Paris: FR0014003TT8, DSY.PA), Medidata is headquartered in New York City and has been recognised as a Leader by Everest Group and IDC. Discover more at  Listen to our latest podcast, from Dreamers to Disruptors, and follow us at @Medidata.

Our team

At the heart of Medidata's ecosystem, the Data Platform Team powers and connects every application, serving as the engine for enterprise data convergence. Every interaction across our global platform generates critical data—and our mission is to transform that raw information into high-impact insights.

Operating on a Data-as-a-Product philosophy, our team builds the foundational infrastructure. This infrastructure includes high-throughput streaming and cloud warehousing to automated data governance. It fuels clinical analytics, AI/ML innovations, and global data sharing.

Why Join Us?

  • Direct Impact at Scale: Promote the central data engine behind every Medidata application, directly accelerating clinical trials and life-saving operational outcomes globally.
  • Modern Distributed Stack: Build at the intersection of real-time event streaming pipelines, scalable cloud data warehouses, and enterprise-grade automated data security.
  • Product-Minded Engineering: Treat data as a first-class product, transforming static databases into high-value, reusable assets for internal AI/ML teams and external partners.
  • Fuel Advanced AI/ML
    :
    Promote next-generation predictive modelling and clinical analytics by standardising and safeguarding complex healthcare datasets.

What will you do

Reporting to a Director of Engineering, as a Principal Data Engineer / Architect
, you will lead the strategic vision and hands-on execution of our next-generation Object-Centric Data Fabric
. You will transition traditional application-centric architectures into a centralised semantic layer that seamlessly unifies multi-stream operational data—including Electronic Data Capture (EDC), patient telemetry, and real-world health datasets. In this role, AI augmentation is natively woven into your workflow. It acts as a force multiplier to automate routine mapping, query optimization, and regulatory documentation. This allows you to focus on driving high-impact platform architecture.

  • Data Fabric & Lake Architecture: Architect and evolve the enterprise semantic data fabric, converting multi-stream clinical execution datasets into an object-centric model. Design and execute a modern Data Lake strategy centred on Apache Iceberg as the core storage format, ensuring high-performance querying and seamless interoperability with Snowflake and heterogeneous compute engines.
  • AI-Accelerated Schema & Pipeline Engineering: Develop and maintain end-to-end multi-stream ingestion pipelines for complex clinical trial schemas. Use AI-driven schema inference and ontology alignment tools to auto-draft mapping artifacts, dramatically reducing integration timelines across different life science datasets.
  • High-Throughput Streaming & Backend Services: Build scale, fault-tolerant real-time ingestion pipelines using Kafka, AWS, and Snowflake. Write robust enterprise services in Java or Scala, leveraging AI coding assistants for rapid code generation, refactoring, and performance tuning.
  • Technical Strategy & Database Optimization: Promote technical direction and engineering best practices across teams for Change Data Capture (CDC), clustering, data migration, and aggregation. Use AI query-optimization tools to analyse execution plans,…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary