Data Engineer - Remote2356992 | Minnetonka, Minnesota | Remote
Minnetonka, Hennepin County, Minnesota, 55345, USA
Listed on 2026-08-14
-
IT/Tech
Data Engineering
Data Engineer
At United Healthcare, we're simplifying the health care experience, creating healthier communities and removing barriers to quality care. The work you do here impacts the lives of millions of people for the better. Come build the health care system of tomorrow, making it more responsive, affordable and optimized. Ready to make a difference? Join us to start Caring. Connecting. Growing together.
This role is responsible for designing, developing, and maintaining scalable and reliable data pipelines that support both batch and real-time analytics within an Azure-based data platform. The position operates as part of a collaborative data engineering team, working closely with fellow data engineers and data science & reporting partners to meet evolving data requirements. The scope of the role includes end-to-end pipeline development using Azure Data Factory, Databricks, PySpark, and streaming technologies;
implementation of the Medallion (Bronze/Silver/Gold) architecture; enforcement of data quality, reliability, and performance standards; and adherence to enterprise data governance, security, and documentation practices across the data lifecycle.
You'll enjoy the flexibility to work remotely from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities:
- Pipeline Development:
Design, develop, and maintain robust pipelines to ingest data from various sources (both streaming and batch) into the analytics environment using Azure Data Factory and PySpark via Databricks. Set up real-time data ingestion using tools like Spark Structured Streaming and batch ETL jobs for periodic data loads. Ensure these pipelines are scalable, efficient, and fault-tolerant to handle growing data volumes and velocity - Implement Data as per Medallion Architecture:
Utilize the Medallion (Bronze/Silver/Gold) architecture principles to organize data processing stages. Establish raw data capture (bronze), perform cleansing and transformations (silver), and curate refined datasets for analysis and machine learning (gold). Apply best practices in each layer, such as schema enforcement and checkpointing for streaming data - Optimize Spark jobs by tuning configurations, improving query logic, and managing resources to achieve high throughput and low latency. Address bottlenecks in streaming pipelines (e.g., by scaling clusters or tweaking batch intervals) and ensure timely data delivery. Optimize job scheduling and cluster utilization to balance timely data delivery with cost-effectiveness
- ETL Development & Maintenance:
Build and maintain data pipelines with an emphasis on data cleaning steps. Integrate data from various sources (APIs, databases, file feeds, IoT streams, etc.) into the data platform, writing transformations that handle anomalies (e.g., missing or corrupt values) and standardize datasets. Collaborate with the other data engineer to share responsibility across different pipelines or sources, ensuring redundancy and knowledge transfer - Data Quality Management:
Implement comprehensive data validation rules and checks within pipelines. For example, verify schema correctness, check value ranges for sensor or health data, and ensure referential integrity where applicable. Set up automated alerts or logs that flag inconsistent or bad data, enabling quick intervention. Over time, build a library of data quality tests that run as part of the pipeline (for both streaming and batch processes) to catch issues early - Emerging Pipeline Frameworks:
Leverage modern pipeline frameworks and tools to improve development productivity. For example, use Databricks Delta Live Tables or Lakehouse pipelines to declaratively define data flows where applicable. Explore the use of Spark Declarative Lakeflow Pipelines or similar technologies to simplify the orchestration of complex data processes - Reliability &
Collaboration:
Implement monitoring and alerting for pipeline health. Investigate and resolve problems such as data delays, pipeline failures, or data inconsistencies. Use…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).