×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer (SRE

Job in San Francisco, San Francisco County, California, 94102, USA
Listing for: Thinking Machines Lab
Full Time position
Listed on 2026-08-10
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 350000 - 475000 USD Yearly USD 350000.00 475000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer (SRE)

Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research teams to make every layer of the system more robust and resilient.

What You'll Do
  • Define and own end-to-end reliability, from CI/CD flows to production observability and incident response.
  • Develop appropriate Service Level Objectives for distributed training systems, balancing job completion reliability and scheduling latency with development velocity.
  • Design and implement monitoring and observability across the full training path.
  • Drive incident response for Tinker platform issues, ensuring rapid recovery, thorough incident reviews, and systematic improvements that prevent recurrence.
  • Harden multi-tenant isolation and resource scheduling so that LoRA-based workload co-scheduling maximizes utilization without compromising reliability or data separation
  • Collaborate with security teams to address production vulnerabilities
  • Skills and Qualifications

    Minimum qualifications:

    • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
    • Experience in distributed systems, cloud infrastructure, or site reliability engineering.
    • Proficiency writing software to solve reliability problems, including building tooling and automation.
    • Experience with production incident response, postmortems, and systematic reliability improvement.
    • Strong communication skills and track record of coordination across engineering and research teams.

    Preferred qualifications — we encourage you to apply if you meet some but not all of these:

    • Deep experience operating production cloud services at scale (e.g., public cloud platforms, internal cloud services)
    • Background in distributed training frameworks and how infrastructure failures surface in training behavior.
    • Track record building checkpoint and recovery systems for long-running distributed jobs.
    • Expertise in Kubernetes at scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads.
    Logistics
    • Location:

      This role is based in San Francisco, California.
    • Compensation:
      Depending on background, skills and experience, the expected annual salary range for this position is $350,000 – $475,000 USD.
    • Visa sponsorship:
      We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
    • Benefits:
      Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
    To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
    (If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary