×
Register Here to Apply for Jobs or Post Jobs. X

Platform Infrastructure Engineer (SRE Core

Job in Gatineau, Province de Québec, G8R, Canada
Listing for: Menlo Security
Full Time position
Listed on 2026-08-03
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
Job Description & How to Apply Below
Position: Platform Infrastructure Engineer (SRE Core)

Menlo Security's mission is enabling the world to connect, communicate and collaborate securely without compromise. COVID-19 has made our mission all the more real. We support customers across various enterprises including Fortune 500 companies, 9/10 of the largest global banks and the Department of Defense.

The world has fundamentally changed. We are growing from 400 employees into the next phase of our journey, and we need passionate talent filled with empathy and agility. The right candidate for the job is ethical, hyper-organized, fanatical about seeing things through to completion, service-oriented, and humble enough to take feedback and coaching yet confident enough to provide feedback and coaching.

Menlo is well-funded for growth and our investors are second to none. They include Vista Equity Partners (Vista), General Catalyst, JPMC, American Express, HSBC, and Ericsson Ventures.

Summary

Platform Infrastructure Engineering builds and operates Menlo Security's Infrastructure Platform, enabling our customers to connect to the Internet without compromise. As a Platform Infrastructure Engineer, you'll join a globally distributed team of experienced engineers building and managing the company's core infrastructure services on a cloud-native platform built on Google Kubernetes Engine and VMs spanning multiple regions and environments. The team manages infrastructure as code with Terraform and Spacelift, deploys with Helm, and emphasizes security-first design, comprehensive observability, and multi-region resilience.

The team also uses AI-assisted development and code-review tools, including Gemini Code Assist, as part of the standard engineering workflow, and this role is expected to use LLM-based tooling to build and troubleshoot infrastructure code efficiently.

Outcomes & KPIs

Key Outcome(s) Owned:

  • Reliable, secure, and scalable infrastructure across GCP and AWS supporting Menlo's platform globally.

  • Reduced operational toil and incident recurrence through automation and Infrastructure as Code practices.

  • Comprehensive, end-to-end observability framework providing deep platform visibility, proactive health monitoring, and accelerated incident detection and resolution.

Success Metrics / KPIs:

  • Infrastructure uptime/availability across regions (e.g., 99.9%+)

  • Mean time to detect (MTTD) and mean time to resolve (MTTR) for incidents

  • Percentage of infrastructure changes deployed via IaC (Terraform) vs. manual changes

  • On-call incident volume and reduction in repeat/preventable incidents

  • Lead time for provisioning new infrastructure

What You'll Do
  • Implement, deploy, and maintain VM and Kubernetes infrastructure on GCP and AWS across dozens of clusters spanning development, staging, and production environments in multiple regions

  • Build and maintain Infrastructure as Code using Terraform modules and Spacelift (or equivalent TACOS), provisioning networking, compute, storage, and security components, and implementing multi-layer configuration management workflows

  • Implement and maintain observability solutions using Grafana Cloud, Prometheus/Mimir, and OTel collectors, designing dashboards and alerting rules across all platform components

  • Manage certificate lifecycle, DNS automation, ingress controllers, and service mesh networking with Cilium

  • Partner with peers and across Engineering, Product, Compliance, and Security teams to align on requirements and consult on capacity planning, disaster recovery, and architectural decisions

  • Identify and eliminate toil through automation — writing scripts, building CI/CD pipelines, and using AI-assisted coding tools to move faster

  • Participate in a 24x7 on-call rotation as part of a globally distributed team, responding to incidents and driving post-incident reviews

  • Functional Competencies

    Required:

    • Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience

    • Proficiency in common programming and scripting languages, particularly Python, Bash, and Go

    • Understanding of network topologies, communication protocols (e.g., TCP/IP, HTTP/S, UDP, TLS), and enterprise-grade connectivity solutions

    • Kubernetes expertise, including cluster administration, RBAC,…

    Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
    To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary