Principal Data Engineer – Databricks | Spark | Delta Lake | PySpark | Data Lakehouse | AWS/Azure
Job Description
Location: Toronto
Work Model: Onsite (4 days/week)
Key Requirements
• 12–18 years of overall Data Engineering experience.
• 8+ years of experience with Enterprise Data Warehouse and Data Lake platforms.
• 5+ years of hands-on experience with Databricks and Apache Spark at scale.
• Strong experience modernizing legacy Cloudera platforms (CDH/CDP, Hive, HBase, Impala, Spark) to Databricks Lakehouse.
• Experience redesigning ingestion, transformation, and consumption patterns from HDFS-based architecture to cloud object storage and Delta Lake.
• Experience refactoring legacy Hive/Impala logic into PySpark and Spark SQL ELT pipelines.
• Experience ensuring data reconciliation, audit integrity, and consistency during migration.
• Experience designing and governing Enterprise Data Warehouse and Data Lake/Lakehouse architectures.
• Experience implementing layered architectures including:
• Raw/Landing Layer
• Curated/Conformed Layer
• Semantic/Consumption Layer
• Experience modernizing traditional Enterprise Data Warehouse platforms into scalable Lakehouse architectures.
• Strong experience with finance and risk data models, including:
• General Ledger
• Sub-ledger
• Financial Hierarchies
• Credit Risk Models
• Liquidity Risk Models
• Market Risk Models
• Experience enabling reporting use cases including aggregation, drill-down, and drill-back capabilities.
• Experience building and managing semantic/consumption layers for BI, reporting, and analytics.
• Ability to define business metrics, dimensions, hierarchies, and KPIs.
• Experience with Databricks SQL, Delta Tables, and dbt or similar frameworks.
• Strong experience developing and optimizing large-scale data pipelines using:
• Py Spark
• Spark SQL
• Delta Lake
• Experience implementing Medallion Architecture:
• Bronze Layer
• Silver Layer
• Gold Layer
• Experience optimizing workloads using Z-ORDER, OPTIMIZE, caching, and cluster configurations.
• Experience implementing data governance, data quality frameworks, reconciliation controls, and exception handling.
• Experience establishing data lineage and metadata management.
• Knowledge of data security, access control, and compliance standards.
• Experience with cloud platforms such as AWS or Azure.
• Experience with CI/CD pipelines using:
• Git
• Terraform
• Jenkins
• Azure Dev Ops
• Familiarity with orchestration tools such as:
• Apache Airflow
• Databricks Workflows
• Experience with dbt is a plus.
• Ability to act as a technical authority and lead architecture decisions.
• Experience mentoring senior engineers and establishing engineering standards.
• Strong stakeholder management skills with finance, risk, analytics, and governance teams.
• Ability to translate complex data structures into business-ready insights.
Nice to Have
• Experience in Banking, Financial Services, Insurance (BFSI), Capital Markets, or regulatory reporting.
• Exposure to:
• SAP Finance
• Oracle Financials
• SAP S/4
HANA
• Experience supporting AI/ML workloads.
• Databricks or cloud certifications.
Key Responsibilities
• Lead Cloudera to Databricks transformation initiatives.
• Design and implement enterprise Data Lakehouse and Data Warehouse solutions.
• Build scalable, high-performance data pipelines and modern data architectures.
• Drive data modernization, governance, quality, and security initiatives.
• Support regulatory, management, and analytical reporting platforms.
• Provide technical leadership, mentor engineering teams, and establish best practices.
RequirementsSailpoint
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: