Data Engineer
Job in
Indianapolis, Hamilton County, Indiana, 46262, USA
Listed on 2026-08-30
Listing for:
Scorpion Therapeutics
Full Time
position Listed on 2026-08-30
Job specializations:
-
IT/Tech
Data Engineering, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
What You Will DoData Engineering & Pipeline Development
- Design, develop, and optimize scalable data pipelines using Databricks, PySpark, Python, SQL, and Delta Lake to ingest, transform, and load data into data warehouses and data lakes.
- Build Databricks pipelines implementing canonical data models across the medallion architecture (Bronze → Silver → Gold) within the CE trust boundary.
- Evaluate/apply Databricks capabilities (Unity Catalog, Delta Lake, Databricks Workflows, serverless compute, Lakebase, ingestion connectors) based on performance, cost, and scalability.
- Implement and maintain ELT/ETL workflows (Databricks Workflows, Auto Loader, Structured Streaming, Delta Live Tables).
- Build and maintain CI/CD pipelines (Git Hub Actions; dev → test → prod promotion) for CE data and artifacts.
- Automate data ingestion and product creation to reduce manual maintenance and onboarding.
- Implement data governance policies ensuring data quality, integrity, security, and compliance (e.g., GxP, HIPAA) and covered-entity constructs.
- Implement row/column-level security, masking, and tokenization to enforce PHI isolation.
- Use Unity Catalog for metadata, lineage, and access control.
- Establish testing/validation (pytest, DLT/Great Expectations) and monitoring/alerting.
- Develop/maintain data models, schemas, metadata; follow Lakehouse/Medallion principles.
- Create reusable transformation frameworks and automated data quality checks.
- Partner on reference architecture; document pipelines/processes.
- Translate stakeholder data requirements into technical solutions.
- Monitor performance, troubleshoot, and improve availability/reliability.
- Participate in code reviews/architecture discussions; promote best practices.
- Evaluate/recommend new tools; support production integration of ML/AI tools with PHI classification/consent.
- Participate in Agile ceremonies (Jira or equivalent).
- Bachelor’s degree in CS/Engineering/IS or related quantitative field.
- 5+ years data engineering/ETL development.
- Proficient in SQL and at least one language (e.g., Python/Java/Databricks).
- Experience with cloud data platforms (Databricks, AWS/Azure/GCP) and services (S3, Redshift, Snowflake, ADLS, Big Query).
- Proficient with Git-based CI/CD workflows (Git Hub Actions or equivalent).
- Excellent problem-solving; ability to build/test pipelines from architecture.
- Strong communication/collaboration.
- Prior pharma/life sciences experience (preferred).
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×