More jobs:
Job Description & How to Apply Below
We are looking for a Senior, Super Hands-On Databricks Data Engineer who lives and breathes code, query optimization, and modern data architecture. In this role, you won't just design architectures onwhiteboards—you will write production PySpark/SQL, optimize Databricksclusters, build streaming and batch pipelines, and enforce data governance.
You will own end-to-end pipeline execution from rawingestion to curated Gold layer models, playing a lead role in modernizing our Lakehouse platform.
Key Responsibilities
1. Hands-On Pipeline Development & Lakehouse Architecture
Design,build, and maintain enterprise-scale batch and real-time streamingpipelines using PySpark, SQL, Delta Live Tables (DLT), and Auto Loader .
Implement and refine Medallion Architecture (Bronze Silver Gold) to support downstream BI,reporting, and Machine Learning workloads.
Enforceschema evolution, ACID transactions, and data compaction using Delta Lake core constructs .
2. Performance Tuning & Optimization (Deep Tech)
Diagnose and resolve Spark performance bottlenecks: data skew, OOM errors,excessive shufflings, and memory spills .
Optimizequeries using Liquid Clustering, Z-Ordering, Data Partitioning, AQE(Adaptive Query Execution), and Photon engine tuning .
Benchmark and optimize Databricks compute workloads to minimize DBU (Databricks Unit) consumption and cloud costs (Fin Ops) .
3. Governance, Security & Quality
Implementend-to-end data governance, fine-grained access control (row/column-level security), and lineage tracking using Unity Catalog .
Automateautomated data quality validation checks and alert mechanisms across the pipeline life cycle.
4. Operations, CI/CD & Dev Ops
Automatepipeline orchestration using Databricks Asset Bundles (DABs) or Databricks Workflows / Apache Airflow .
BuildCI/CD pipelines (Git Hub Actions, Azure Dev Ops, or Git Lab) for automated testing, deployment, and code promotions.
Requirements
Required Skills &Qualifications
Must-Haves
Experience:
8+ years in Data Engineering , with 4+ years of intensive, hands‑on production experience on Databricks .
Programming Mastery: Fluent in PySpark, Advanced SQL , and Python.
Databricks Ecosystem: Deep experience with Delta Lake, Unity Catalog, Delta Live Tables (DLT), Auto Loader, and Databricks Workflows .
Cloud
Infrastructure: Strong hands‑on experience in at least one primarycloud provider ( AWS, Azure, or GCP ) integration with Databricks(S3/ADLS Gen2, IAM, Key Vaults/Secret Manager).
Data Modeling: Solid understanding of dimensional modeling (Kimball), One Big Table (OBT) strategies, and data vault patterns.
CI/CD& Software Engineering: Proficient in Git workflows, unit testing
PySpark code (pytest), and deployment automation.
Preferred / Nice-to-Haves
Certifications:
Databricks Certified Data Engineer Professional.
Streaming: Hands‑on with Apache Kafka, Event Hubs, or Kinesis integration via Structured Streaming.
GenAI/ ML Ops: Familiarity with MLflow, Feature Store, or Vector Searchwithin Databricks.
Infrastructure as Code (IaC): Experience using Terraform to provision Databricksworkspaces and storage resources.
Performance Indicators(How success is measured)
Pipeline Reliability: Maintaining strict SLA thresholds on critical Gold-layermodels.
Cost Efficiency: Measurable reduction in DBU costs through effectivecompute profiling and tuning.
Code Quality: High test coverage and zero-downtime CI/CD deployments.
#J-18808-Ljbffr
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×