More jobs:
Data Engineer; Databricks
Job in
Ridley Park, Delaware County, Pennsylvania, 19078, USA
Listed on 2026-07-26
Listing for:
Diverse Lynx
Full Time
position Listed on 2026-07-26
Job specializations:
-
IT/Tech
Data Engineering
Job Description & How to Apply Below
Data Engineer (Databricks)
Location:
Seattle, WA St Louis, MO/ Dallas, TX/ Plano, TX/ Charleston, SC/ Ridley Park, PA (Onsite)
Experience
Required:
8 - 10 years of experience
Pay Range: $55 to $65/Hr On W2
Must Have Technical/Functional Skills- Successfully executed a data migration or modernization to Databricks, preferably IBM Data Stage to Databricks on AWS
- Experience in handling large migrations to Databricks
- Good analytical skills to compare the legacy and modern data platform end to end right from source to target
- Good understanding of Databricks implementation of Medallion layer architecture
- Independently lead and managed large Databricks migrations
- CI/CD Integration:
Implement version control (e.g., Git) and automated deployment processes for Databricks assets
- Experience in advanced SQL for building modular analytics workflows, utilizing advanced Common Table Expressions (CTEs), and writing high-performance queries inside Databricks SQL Analytics
- Experience in Python or Scala to build, optimize, and debug complex data transformation scripts, custom functions, and machine learning pipelines
- Experience in Apache Spark ecosystem for understanding cluster execution flow, memory allocation, driver/worker nodes, and handling data frames
- Experience in Delta Lake architecture to understand ACID transactions on object storage, data skipping, partition strategies, and automated data compaction
- Experience in Delta Live Tables (DLT) & Workflows for constructing and orchestrating production-ready, declarative streaming, and batch ETL pipelines
- Experience in Unity Catalog for setting up data governance, column/row-level access control, and tracking end-to-end data lineage across work spaces
- Experience in Auto Loader for implementing modern, incremental data ingestion patterns from cloud blob storage into the lakehouse
- Pipeline Conversion:
Translate visual Data Stage Parallel Jobs and Sequences into Python/PySpark scripts or Databricks Notebooks - Legacy Refactoring:
Modernize legacy logic rather than applying "lift and shift" anti-patterns; adapt workflows to think in distributed Data Frames rather than Data Stage stages - Logic Mapping:
Map Data Stage components—such as Aggregators, Joiners, Transformers, and Sort stages—to equivalent Spark operations
- Validation & Reconciliation:
Build automated reconciliation frameworks to compare row counts, checksums, and aggregate sums between legacy Data Stage outputs and new Databricks output - Data Cleansing:
Identify and resolve data type discrepancies, null-handling differences, and encoding issues during the extraction and loading phases
- Orchestration:
Replace Data Stage sequence jobs with Databricks workflows (or external orchestrators like Azure Data Factory/Airflow) to schedule and manage dependencies - Data Governance:
Enforce data lineage, security, and cataloging using Unity Catalog to ensure compliance in the new Lakehouse environment
- Cloud Providers (AWS):
Understanding underlying cloud object storage, identity access management (IAM), and network security configurations - Dev Ops & Bundles:
Familiarity with Databricks Asset Bundles (DABs) and CI/CD tools to automate the deployment of work spaces and pipeline assets
- Code Conversion & Translation:
The ability to parse legacy code structures and refactor them into Databricks-native code
Skills in using AI coding assistants and open framework agent tools to analyze application interdependencies, automate schema mapping, and accelerate lift-and-shift workloads
The ability to parse legacy code structures from ETL pipelines, Informatica, data Stage preferred Experience working in Agile teams and understanding of data governance frameworks.
ResponsibilitiesSupport post-migration environment from IBM Data Stage to Databricks Incident & Lifecycle Management
- CI/CD Deployment:
Support code deployments across Development, Test, and Production environments using Databricks Repos and REST APIs - Monitoring & Alerting:
Set up monitoring via Databricks System Tables and observability tools to catch job failures, data anomalies, or latency spikes early
Pipeline Maintenance & Orchestration
- Workflow Management:
Transition from Data Stage job sequences to native databricks workflows for scheduling, dependency tracking, and alerts - ETL Refactoring:
Troubleshoot and fix issues in generated PySpark or Spark SQL code that replaced legacy Data Stage Transformer or Lookup stages - Streaming & Batch Integration:
Support ongoing data ingestion using databricks autoloader to process files continuously from cloud storage
Performance Tuning & Cost Optimization
- Compute Management:
Monitor and configure serverless or classic clusters to prevent over-provisioning - Query…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×