Senior Data Engineer
Listed on 2026-09-30
-
Software Development
Data Engineering
About Us
A pioneer in the caregiving space, Careforth supports family caregivers across the United States to confidently care for their loved ones ough a combination of in-person home visits, remote coaching and our proprietary digital collaboration app, we provide caregivers with support, guidance, confidence, and connection to resources they need. The Caregivers and families we support stay with Careforth for many years, building lasting relationships along the way.
Join us today and live our values: lead with heart, cultivate trust, go beyond.
The Senior Data Engineer will be a cornerstone of Careforth’s data platform team, architecting and operating the cloud-native, lakehouse-based infrastructure that powers our full suite of AI and analytics products — from clinical and caregiver risk intelligence to LLM-enabled chatbots, real-time signal capture, and payer-facing outcome reporting.
The ideal candidate brings deep, hands‑on expertise with Databricks, Apache Spark, and AWS Cloud services, and can serve as both a technical leader and mentor. You will own end‑to‑end pipeline design from raw source ingestion through curated analytics‑ready datasets, ensuring data systems are reliable, scalable, and HIPAA‑compliant.
What You Will Do- Architect and build robust, scalable, and secure data pipelines leveraging Databricks (Delta Lake, Spark, Unity Catalog, Workflows) and AWS (S3, Redshift, Athena, Glue, Lambda, Step Functions, Kinesis), supporting both batch and real‑time processing modes.
- Design and evolve the enterprise lakehouse architecture, including data modeling, Delta table optimization, data versioning, and a semantic data layer and metrics store serving BI tools and downstream ML feature engineering.
- Build and maintain ETL/ELT orchestration (Airflow, dbt) and stream processing pipelines (Kafka/Kinesis) ingesting claims, EHR/EMR, pharmacy, IoT, survey, CRM, and digital engagement event data.
- Partner with Data Scientists to design and maintain feature store architecture, real‑time member context assembly, and temporal feature computation services supporting risk and predictive models.
- Develop data infrastructure for LLM and NLP use cases: vector databases, knowledge base ingestion, transcript storage and indexing, secure audio storage, and LLM response logging and audit trails.
- Build APIs and integrations for downstream consumers: care team dashboards, lead scoring, pathway recommendation, queue prioritization, CRM connectors (Salesforce, Hub Spot), and payer‑facing data exchange.
- Implement HIPAA‑compliant data handling across all pipelines including PHI detection, redaction, audit trails, row‑level security, SSO, and secrets management (AWS Secrets Manager, Vault).
- Build reporting infrastructure including a report template engine, multi‑format export (PDF, Excel, CSV), scheduled report generation, and delivery APIs for internal and payer‑facing reporting.
- Conduct pull request reviews, enforce engineering best practices, mentor junior engineers, and research and apply emerging data engineering tools and patterns.
- Participate in design discussions with technical leads across product lines; perform other duties and special projects as assigned.
- Bachelor’s or Master’s Degree in Computer Science, Data Engineering, or a related field.
- 7–10 years of professional data engineering experience, including 3+ years with modern cloud‑based data lake or lakehouse architectures.
- Deep expertise in Apache Spark (PySpark, Scala, or Java) for large‑scale distributed data processing and strong hands‑on Databricks experience across Delta Lake, Unity Catalog, and Workflows.
- Hands‑on experience with AWS data services: S3, Redshift, Athena, Glue, Lambda, Step…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).