Principal Data Engineer
Listed on 2026-07-22
-
Software Development
Data Engineering
Location: Philadelphia, PA
Onsite Flexibility: Hybrid
Job Details- Position Type: Right to Hire
- Contract Duration: 6 month conversion
- Pay Rate: $85.00 $ 90.00 / Hour (USD)
- Schedule: 2 days on-site Tuesday / Wednesday
- Work Authorization: Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time.
This role is focused on building and leading the data engineering foundation that powers real-time decisioning, operational applications, analytics, ML/AI model development, and data services across a medical client. The Principal Data Engineer will own the design, delivery, and maturity of production-grade data pipelines and data platforms, with a primary emphasis on real-time streaming, IoT telemetry, Databricks, Azure, data services for APIs and microservices, and reliable data products for downstream consumption.
We are looking for a Principal Data Engineer to serve as a hands-on technical and people leader for data engineering, data platform architecture, real-time streaming, and production data services. This role will focus on designing, building, operating, and improving data pipelines and data products while also bringing principal-level judgment to architecture, stakeholder shaping, delivery priorities, team management, and production readiness.
This is a hands-on engineering leadership role first. The ideal candidate should be comfortable spending significant time working directly with Databricks, Spark, SQL, Python/PySpark, Azure services, streaming architectures, data quality frameworks, pipeline automation, CI/CD, and production troubleshooting. They should also be able to operate with the maturity of a principal-level leader: shaping unclear requirements, making pragmatic technical decisions, managing and mentoring engineers, and driving work forward without waiting for perfect specifications.
This is a fast-moving, startup-like environment. Requirements may be incomplete, priorities may evolve, and the right candidate will help create clarity while building quickly. We need someone who can move from ambiguous business need to reliable data capability with urgency, discipline, and ownership.
Stakeholder shaping is a critical part of this role. The Principal Data Engineer should be able to work directly with business, product, software engineering, analytics, ML/AI, operations, and leadership stakeholders to define what data needs to exist, how it should be consumed, what production guarantees are required, and how success should be measured.
A background in commercial software, SaaS, digital products, healthtech, fintech, IoT, data platforms, or other product-driven environments is strongly preferred. We want someone who understands that data pipelines and data services are not just technical artifacts. They are product capabilities that support real users, real workflows, operational decisions, ML/AI systems, APIs, analytics, and measurable business outcomes.
Key ResponsibilitiesHands-On Data Engineering and Platform Development
- Design, build, optimize, and operate production-grade batch and streaming data pipelines on Azure and Databricks, with a primary focus on real-time IoT and telemetry use cases within a Medallion architecture.
- Develop ETL/ELT workflows to ingest, transform, validate, and serve large volumes of structured, semi-structured, unstructured, and streaming data.
- Build and maintain reliable data products, data services, APIs, and microservices that support operational applications, analytics, software engineering, and ML/AI teams.
- Use Python, PySpark, Spark SQL, SQL, Delta Lake, Databricks Workflows, CI/CD, and related tools to build maintainable, testable, and observable data systems.
- Troubleshoot complex production pipeline issues across Databricks, Azure, streaming systems, APIs, and source systems, including root cause analysis, corrective action, and prevention planning.
- Move quickly from rough business need to prototype, pilot, and production-ready data capability while maintaining appropriate engineering discipline.
Real-Time Streaming, IoT Telemetry, and Operational Data Services
- Lead the design and delivery of real-time streaming ingestion and processing patterns for connected medical device telemetry, event data, and operational data feeds.
- Implement streaming solutions using Azure Event Hubs, Azure Stream Analytics, Databricks, Delta Lake, and related Azure integration patterns.
- Design cost-effective throughput, partitioning, delivery, retention, and replay strategies for high-volume event and telemetry workloads.
- Create consumption patterns that support APIs, microservices, operational applications, near-real-time decisioning, analytics, and ML/AI use cases.
- Define reliability, latency, quality, observability, and supportability expectations for production streaming systems.
Databricks, Lakehouse, and Data Platform Architecture
- Set direction for Databricks-based data…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).