×
Register Here to Apply for Jobs or Post Jobs. X

Head of Engineering; Infrastructure & Site Reliability Engineering; SRE

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Postman
Full Time position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, Network Engineer
Salary/Wage Range or Industry Benchmark: 260000 - 380000 USD Yearly USD 260000.00 380000.00 YEAR
Job Description & How to Apply Below
Position: Head of Engineering (Infrastructure & Site Reliability Engineering (SRE)
  • Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence
  • As Head of Infrastructure, you’ll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement
  • You’ll own the infrastructure that underpins one of the world’s most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution
  • In addition to infrastructure, you’ll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale
  • You’ll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization
  • Hire, manage, mentor, and coach a geographically distributed team of infrastructure and SRE engineers across SF Bay Area, India, and Europe, helping them grow both technically and professionally
  • Drive and uphold a culture of respect, integrity, inclusion, ownership, and accountability within the team
  • Set clear goals and provide regular feedback to ensure your team is motivated and aligned with the platform and company vision
  • Build a high-performing team with the skills and practices to reliably operate and evolve cloud agnostic infrastructure at scale
  • Foster psychological safety and cross-regional collaboration across time zones
  • Own the architecture and evolution of Postman’s cloud agnostic infrastructure, driving improvements that increase reliability, performance, and cost efficiency at scale
  • Lead the design and implementation of infrastructure improvements across Kubernetes, Cluster API, Argo, Helm, Crossplane, service mesh (Istio), AWS, and Azure environments
  • Own the SRE function end to end: SLIs/SLOs, error budgets, capacity planning, incident management, and reliability engineering practices across the platform
  • Set the technical direction for how Postman’s infrastructure and reliability practices evolve to support a large engineering organization with hundreds of services and dozens of teams, with a focus on enabling product teams to operate autonomously
  • Partner with platform, security, and product engineering teams to ensure infrastructure and reliability decisions align with broader company goals
  • Ensure infrastructure and reliability best practices are upheld across the organization, including Git Ops, CI/CD, observability, on-call, and incident response
  • Own the infrastructure and reliability roadmap, balancing operational reliability with longer-term architectural investments
  • Break down complex infrastructure and reliability initiatives into clear, actionable milestones and manage delivery on time and at high quality
  • Proactively identify and resolve roadblocks, working across teams to unblock engineering work and minimize customer impact
  • Work closely with stakeholders across engineering, product, and security teams to align on infrastructure and reliability priorities and constraints
  • Foster open communication within and across teams, promoting transparency on system health, risk, error budgets, and roadmap
  • Represent infrastructure and SRE in leadership forums, advocating for technical needs and communicating clearly on trade-offs
  • Own Postman’s SRE function, establishing and evolving on-call practices, escalation policies, monitoring, and incident management across the company
  • Define and track SLIs/SLOs and error budgets for critical services, using them to guide investment decisions and prioritization
  • Drive a blameless postmortem culture, ensuring incidents produce durable learnings, clear action items, and measurable follow-through
  • Drive a culture of continuous improvement, learning from incidents, automating toil, and reducing operational burden so that product engineering teams can ship independently
  • Champio…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary