×
Register Here to Apply for Jobs or Post Jobs. X

Staff Field Reliability Engineer

Job in San Diego, San Diego County, California, 92189, USA
Listing for: honeycomb.io
Full Time position
Listed on 2026-08-15
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 200000 - 240000 USD Yearly USD 200000.00 240000.00 YEAR
Job Description & How to Apply Below

What We’re Building

Honeycomb is a service for the near and present future, defining observability and raising expectations of what developer tools can do! We’re working with well known companies like Hello Fresh, Slack, Launch Darkly, and Vanguard and more across a range of industries. This is an exciting time in our trajectory, we’ve closed Series D funding, scaled past the 200-person mark, and were named to Forbes’ America’s Best Startups of 2022 and 2023!

What

We’re Building

Honeycomb is a service for the near and present future, defining observability and raising expectations of what developer tools can do! We’re working with well known companies like Hello Fresh, Slack, Launch Darkly, and Vanguard and more across a range of industries. This is an exciting time in our trajectory, we’ve closed Series D funding, scaled past the 200-person mark, and were named to Forbes’ America’s Best Startups of 2022 and 2023!

If you want to see what we’ve been up to, please check out these blog posts and Honeycomb.io press releases.

Who We Are

We come for the impact, and stay for the culture! We’re a talented, opinionated, passionate, fiercely inclusive, and responsible group of bees. We have conviction and we strive to live our values every day. We want our people to do what they truly love amongst a team of highly talented (but humble) peers.

How We Work

We are a fully distributed company, which means we believe it is not where you sit, but how you deliver that matters most. We invest in our people and care about how you orient to our culture and processes. At the same time we imbue a lot of trust, autonomy, and accountability from Day 1.

Little More About The Team

The Field Reliability Engineer (FRE) is an extension of Honeycomb’s extensive brand and technical expertise focused on our customer base via engagements, support and our managed services. In this role you parachute into the most complex, highest-stakes technical situations our customers face - unblocking them when they’re stuck, creating tooling, guiding our Solution Architecture team through complex decisions, and ensuring prospects and customers get the most out of Honeycomb and the broader observability ecosystem.

You’re equal parts platform engineer and customer engineer. You build and operate the managed infrastructure our largest customers depend on, and you’re the technical backstop the SA team calls when a deal gets deep into infrastructure, data pipelines, or production debugging.

This isn’t a support role - it’s a technical leadership role where you solve problems that don’t have runbooks yet, build the platforms and tooling so the next person can have a runbook, and directly impact revenue by unblocking our most strategic deals.

What You’ll Do Platform Engineering — Architecture & Standards Ownership
  • Define the architecture and operational standards for Refinery as a Service (RaaS) and Honeycomb Private Cloud (HnyPC) — decisions other engineers build within — across multiple AWS accounts and regions.
  • Architect the Terraform modules, Helm charts, and deployment automation that other FREs build on and extend, not just consume.
  • Set the technical direction for how Honeycomb instruments, monitors, and operates its own managed infrastructure — using Honeycomb to monitor Honeycomb.
  • Own capacity planning, scaling strategy, upgrade sequencing, and cost optimization across multi-region AWS environments.
  • Build platforms and automation that change how the FRE team operates at scale — enabling the team to grow without proportional headcount.
Technical Escalation & Unblocking
  • Serve as the final technical escalation point for the most novel, highest-stakes customer situations — problems with no precedent in existing runbooks.
  • Resolve deep infrastructure and observability issues spanning distributed systems, Kubernetes clusters, AWS networking (ALBs, Private Link, NLBs, VPCs), and polyglot service meshes, often in real time under revenue-critical pressure.
  • Partner directly with customer SRE, platform, and engineering leadership to navigate multi-week escalations and architecture redesigns tied to the company's largest relationships.
  • Provide senior incident command…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary