×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer; SRE

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Retool Inc.
Full Time position
Listed on 2026-07-29
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Project Manager, Systems Engineer
Salary/Wage Range or Industry Benchmark: 163710 - 306000 USD Yearly USD 163710.00 306000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer (SRE)

ABOUT RETOOL

Nearly every company in the world runs on custom software for critical operations like tracking performance metrics, handling support workflows, building admin dashboards, and countless processes you might never have thought of. But most companies don't have the resources to properly invest in these tools, leading to a lot of old, clunky internal software, or worse, teams still stuck in manual and spreadsheet workflows.

AI has changed who gets to build software. The definition of "developer" now includes analysts, operators, and domain experts creating solutions directly—and the tools they reach for are multiplying by the week. That's both an opportunity and a challenge: as more people build with more AI tools, the risk of shipping ungoverned software into production grows just as fast. At Retool, we're building the platform that makes all of it safe to ship.

Build with any AI tool you want, then deploy into one place that connects to your real business data, enforces enterprise policies automatically, and lets teams create once and reuse everywhere with shared, trusted components. The cost of building software has collapsed. The cost of governing it hasn't—and that's the problem we solve. Developers and domain experts have already automated over 100 million hours of work on our platform, freeing them to focus on creative problem-solving and strategic work that drives real business value.

The people closest to the problem can now build the software to solve it, safely, and within enterprise guardrails. Let's build the future together.

WHY WE'RE LOOKING FOR YOU:

Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system. Retool's Core Infrastructure team owns the systems that make this possible:
Retool Cloud, managed single tenant environments, BYOC (bring-your-own-cloud) environments, Kubernetes and Helm deployments, Docker Compose, and the migration paths between them. It is a broad surface area, and it is one of the biggest levers we have for making Retool work for enterprise customers. The work is not clean-room infrastructure. Customers run different clouds, different versions, different deployment models, and different levels of operational maturity.

A bad upgrade experience can leave a customer many versions behind. A manual Terraform run can become the bottleneck during a launch or incident. We are hiring SREs who want to turn that mess into leverage. You will help us reduce customer toil, automate upgrades and infrastructure changes, build reliability tooling across Retool Cloud and customer-owned environments, and make Retool easier to deploy and operate at enterprise scale.

The strongest candidates are comfortable debugging Kubernetes, Terraform, AWS, Postgres, networking, and deployment problems, then stepping back and building the automation or product surface that prevents the same problem from happening again.

What you'll do:
  • Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
  • Build the automation that turns today's manual infrastructure work into repeatable systems:
    Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
  • Improve observability for Retool Cloud, self-hosted customers, and internal operators. We care less about exposing every metric and more about turning health signals into clear status, likely causes, and recommended actions.
  • Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current
  • Help move customers from legacy or less-supported deployment models toward supported paths such as Retool's official deployment paths (Blueprints, Kubernetes, and Helm), with migration flows that are repeatable enough for customers, Support, and TAMs to trust.
  • Partner with product engineers on infrastructure requirements for new Retool products,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary