×
Register Here to Apply for Jobs or Post Jobs. X

Senior Platform Engineering

Job in Toronto, Ontario, M5A, Canada
Listing for: Scotiabank
Full Time position
Listed on 2026-07-20
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Azure, SRE/Site Reliability
Job Description & How to Apply Below

Is this role right for you?

In this role, you will:
  • Guidance and Direction: Provide clear direction to the team, set goals, and keep the team accountable for their deliverables. Align team goals with the overall direction of the Azure & Databricks Platform roadmap and enterprise standards.
  • Technical Oversight: Own the technical direction across Azure and Databricks:
    Azure networking and security architecture (VNets, Private Endpoints, NSGs, route tables, Azure Firewall), Azure Identity & Access Management (RBAC, PIM), and Databricks platform governance (Unity Catalog, workspace configuration, cluster policies). Ensure best practices for reliability, cost, and security are consistently applied.
  • Quality Assurance
    :
    Ensure a high quality of support delivery for platform users; adhere to platform SLAs/SLOs and service objectives
  • Process Improvements: Continually improve platform processes and SOPs for efficiency and automation. Design and develop reusable Terraform modules for Azure native resources and Databricks (clusters, SQL warehouses, Unity Catalog objects), enabling consistent, scalable, and automated deployments via Terraform Cloud/Enterprise and CI/CD.
  • Customer Relations: Build strong relationships with data engineers, analysts, and platform users. Communicate proactively with stakeholders and cross‑functional teams (Platform, Security, Cloud Ops, Networking, Data Governance) to align priorities, manage expectations, and drive adoption of platform standards.
  • Advanced Monitoring and Troubleshooting: Troubleshoot and resolve performance issues across Databricks jobs, clusters, SQL warehouses, and Azure dependencies. Implement Azure Monitor and Log Analytics‑based observability with custom dashboards for cluster/job health, driver/executor metrics, and cost insights. Establish proactive alerting and early issue detection via logs/metrics for Databricks and Azure services.
  • Site Reliability: Analyze, triage, and resolve platform issues promptly to achieve SLOs and platform reliability objectives. Drive error‑budget aware practices, post‑incident reviews, and resilience engineering (e.g., autoscaling, retry/backoff strategies, policy guardrails).
  • Incident Management: Provide support during major incidents, including after‑hours support. Lead incident response, communications to users and stakeholders, and root‑cause analysis with clear action items and follow‑through.
  • Observability Tools Development: Design, build, and deploy logging/monitoring solutions for early detection and actionable insights. Standardize ingestion to Log Analytics from Databricks (audit logs, cluster events, job runs) and key Azure resources; built dashboards and alert rules to reduce MTTR.
  • Release Control Management: Maintain and enhance the Infrastructure & Platform release pipeline using Terraform, Terraform Cloud, Azure Dev Ops and/or Git Hub Actions, with source control in Git Hub/Bitbucket and artifact promotion via ACR/Artifacts. Enforce approvals, change windows, and automated checks to ensure safe, repeatable releases.
  • Client Pipeline Management: Implement CI/CD for infrastructure and analytics workloads using Terraform, Docker, Azure Dev Ops/Git Hub Actions, and Artifact/Container registries.

    Automated Terraform plan/apply, Databricks Bundle releases, policy validation, and security scanning to streamline delivery and ensure compliance.
  • Credential Security: Set up Azure Key Vault and Hashi Corp Vault for secret management; integrate with Databricks secret scopes and workload identities. Enforce least‑privilege access via Azure RBAC and rotate credentials per policy.
  • Vendor and Technical Support Interaction: Partner with Microsoft and Databricks support and product teams to fine‑tune and troubleshoot components, plan upgrades, and adopt new capabilities aligned to roadmap and enterprise controls.
  • Mentorship: Mentor junior engineers in best practices for building, deploying, testing, and supporting services on Azure and Databricks. Promote a culture of automation, documentation, and continuous learning.
  • Do you have the skills that will enable you to succeed in this role? We'd love to work with you if you have:

  • 15+ years of IT…
  • Position Requirements
    10+ Years work experience
    Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
    To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary