×
Register Here to Apply for Jobs or Post Jobs. X

Platform Site Reliability Engineer

Job in San Francisco, San Francisco County, California, 94102, USA
Listing for: Specter
Full Time position
Listed on 2026-08-24
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, AWS
Job Description & How to Apply Below

Platform Site Reliability Engineer

We're hiring a Platform Site Reliability Engineer to own the operational health, reliability, and scalability of the cloud platform behind our connected sensor fleet.

This is a high-ownership role at the intersection of site reliability and platform engineering. You'll operate and improve our Kubernetes-based infrastructure, manage cloud resources through Terraform, strengthen observability and incident response, and build the systems that allow our engineering teams to deploy safely and move quickly.

You'll work primarily across our AWS infrastructure and Kubernetes environments while partnering with application, AI, embedded systems, and fleet teams. You'll help resolve production issues when they occur—and then improve the platform so they are less likely to happen again.

Reactive — Triage & Recovery
  • Debug production issues across Kubernetes clusters, Linux systems, AWS infrastructure, networking, and application workloads.
  • Lead incidents from detection through recovery, coordinating across teams when failures span multiple parts of the system.
  • Participate in an on-call rotation and follow incidents through to durable fixes.
Systems Builder — Close the Loop
  • Build, operate, and improve our Kubernetes platform and the AWS infrastructure supporting it.
  • Manage production infrastructure with Terraform, including reusable modules, automated validation, and safe change workflows.
  • Reduce operational toil through automation while improving deployment tooling, CI/CD, and developer workflows.
Observability Owner — Platform Visibility
  • Design and improve observability into system health, functionality, and performance through logging, metrics, tracing, dashboards, and alerting across Kubernetes workloads and AWS infrastructure.
  • Define meaningful service-level indicators and objectives, and close telemetry gaps before they become incidents.
  • Develop runbooks, incident-response procedures, post-incident reviews, and operational readiness standards.
Qualifications
  • Strong Linux systems knowledge and experience diagnosing production systems.
  • Hands-on experience operating Kubernetes in production, including networking, storage, resource management, upgrades, and troubleshooting.
  • Strong experience using Terraform to manage production cloud infrastructure.
  • Experience with AWS, including IAM, networking, compute, storage, and EKS.
  • Solid networking fundamentals, including DNS, load balancing, firewalls, VPNs, subnets, and routing.
  • Experience building operational tooling and automation using Python, Go, Bash, or a similar language.
  • Strong ownership during incidents and the ability to turn ambiguous failures into lasting improvements.
Nice to Have
  • Experience supporting connected devices, edge computing, or on-premises infrastructure alongside cloud systems.
  • Experience with CI/CD, Git Ops, or Kubernetes multi-cluster environments.
  • Familiarity with cloud and Kubernetes security practices.
  • Experience reading firmware logs or low-level Rust or C code when debugging across the edge-to-cloud boundary.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary