Director of Infrastructure & Site Reliability Engineering
Listed on 2026-07-23
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Responsibilities
- We are seeking an experienced engineering leader to lead Alteryx’s Infrastructure, Site Reliability Engineering (SRE), Observability, and Performance Engineering organizations. In this role, you will define the technical vision and execution strategy for the platforms and operational capabilities that power our cloud services, enabling engineering teams to build, deploy, and operate reliable, secure, and highly scalable products
- You will lead multiple engineering teams responsible for cloud infrastructure, reliability engineering, observability, performance optimization, and operational excellence. This leader will partner closely with Product Engineering, Security, Compliance, and Customer Operations to ensure our platform meets the highest standards for availability, scalability, security, and customer experience
- Define and execute the strategy for Alteryx’s centralized Infrastructure, Site Reliability Engineering (SRE), Observability, and Performance Engineering organizations
- Lead the design, operation, and continuous evolution of cloud infrastructure across AWS and GCP, ensuring scalability, reliability, security, and cost efficiency
- Drive Infrastructure-as-Code adoption and governance through Terraform, establishing consistent platform standards, automation, and operational best practices
- Own the company’s observability strategy by building and operating enterprise-grade telemetry platforms using Datadog and related technologies, enabling actionable insights into system health, performance, and customer experience
- Partner with Security, Compliance, and Engineering teams to meet regulatory and customer requirements, including HIPAA, FedRAMP, SOC 2, and other compliance frameworks
- Establish and continuously improve incident management practices, including operational readiness, on-call excellence, postmortem culture, root cause analysis, and measurable reliability improvements
- Develop proactive reliability programs including capacity planning, resiliency testing, disaster recovery, performance benchmarking, and operational risk management
- Define reliability engineering frameworks that enable product teams to own service health through Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, performance objectives, and operational accountability
- Lead the evolution of centralized platform capabilities that simplify how engineering teams build, deploy, monitor, and operate services at scale
- Partner with engineering leadership to improve developer productivity through platform automation, self-service infrastructure, deployment tooling, and operational best practices
- Build, mentor, and develop high-performing engineering managers and technical leaders while fostering a culture of operational excellence, customer focus, accountability, continuous learning, and innovation
Strong understanding of distributed systems, cloud networking, Kubernetes, container orchestration, CI/CD pipelines, and production operations 10+ years of software engineering, infrastructure, or platform engineering experience, with 5+ years leading multiple engineering teams or managers Experience supporting regulated environments and working with compliance frameworks such as HIPAA, FedRAMP, SOC 2, ISO 27001, or similar Strong experience with Infrastructure-as-Code technologies such as Terraform Proven ability to influence technical strategy and drive alignment across engineering, security, product, and executive stakeholders Proven experience leading Infrastructure, SRE, Platform Engineering, or Cloud Operations organizations supporting large-scale SaaS products Experience building and operating modern observability platforms using Datadog, Open Telemetry, Prometheus, Grafana, or similar technologies Excellent communication skills with the ability to translate technical strategy into business outcomes Demonstrated success implementing SRE practices including SLOs, SLIs, error budgets, incident management, operational reviews, and reliability engineering programs Passion for building high-performing teams and developing engineering leaders Deep expertise operating production environments on AWS and/or GCP Familiarity with software performance engineering, load testing, and large-scale distributed systems optimization Experience supporting data-intensive cloud services Experience leading platform transformations for enterprise SaaS organizations If you meet some of the requirements and you share our values, we encourage you to apply
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).