Lead DevOps & Platform Engineer
Listed on 2026-07-22
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Cybersecurity
Hands-on platform ownership with a path to building and leading the function
About UsWe are a growing Vancouver technology startup building proprietary, AI-enabled software and smart digital solutions. We use AI extensively across product development and engineering, while protecting our source code, confidential information, customer data, brands, and other intellectual property. We are now investing in the cloud platform, delivery systems, and operational discipline required to turn rapid innovation into secure, reliable products.
The OpportunityAs our first Lead Dev Ops & Platform Engineer, you will establish and shape the technical foundation of our software development lifecycle — owning the evolution of our Azure platform and software-delivery foundation. This is a hands‑on role for someone who can move comfortably between architecture, implementation, incident response, security, automation, and technical leadership.
You will inherit Azure-hosted products supported by six international developers and a flexible group of Canada-based contractors. Core capabilities, including Infrastructure as Code, centralized observability, mature Git Hub-based CI/CD, and a sustainable production-support model, still need to be established. You will assess the current state, choose pragmatic tools and practices, and bring the process discipline needed to create a foundation that fosters high-quality digital products.
You'll work closely with our Product Manager and development team to implement the standards, tooling, and processes that will scale as our product offerings grow. There are no direct reports initially, but you will set the technical standard for Dev Ops — providing technical direction across the delivery team and determining how work is assigned and executed once it's been prioritized by the Product Manager — while helping leadership determine future staffing needs.
As the company grows, you will participate in hiring and onboarding additional engineers and may have the opportunity to move into formal people leadership.
- Implement industry-standard Dev Ops & Agile/SAFe Agile standards, tools, and governance frameworks, and lead engineering teams in applying these best practices to the software development lifecycle to meet business requirements and user needs.
- Own the architecture, reliability, security, performance, and cost management of the Microsoft Azure platform across development, test, and production environments.
- Implement repeatable Infrastructure as Code using Terraform and/or Bicep, including reviewed modules, environment separation, secure state, drift awareness, and controlled change practices.
- Standardize Git Hub Actions pipelines, repository governance, branch protections, reusable workflows, release controls, and reliable rollback paths.
- Embed automated testing, code-quality checks, web application vulnerability scanning, monitoring and remediation of CVEs/zero-day threats, secrets protection, and deployment safeguards into the delivery lifecycle.
- Select and implement an observability stack from the ground up, covering logs, metrics, traces, dashboards, actionable alerts, service ownership, and appropriate service-level objectives.
- Using Azure and Entra , establish and support secure cloud identity and access management policies including secrets, network-security, conditional access policies, role-based access, privileged identity management (PIM), and Azure Key Vault.
- Lead production incident response, blameless reviews, runbooks, and continual reliability improvements. Critical incidents have historically been infrequent; initially, this role will be a primary after-hours technical escalation point and will establish a more sustainable future coverage model.
- Define and test recovery objectives, disaster-recovery procedures, and business-continuity plans, while improving Azure cost visibility, budgets, alerts, and optimization.
- Contribute to the implementation and management of ITSM best practices using Jira Service Management for Change, Incidents, Problem, and Service Requests.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: