×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer (SRE

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: LeoLabs, Inc.
Full Time position
Listed on 2026-08-12
Job specializations:
  • IT/Tech
    Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 115000 - 192000 USD Yearly USD 115000.00 192000.00 YEAR
Job Description & How to Apply Below
Position: Senior Site Reliability Engineer (SRE)

At Leo Labs, we’re building the living map of activity in space. Through our proprietary global radar network and AI-enabled analytics platform, we collect millions of measurements daily on more than 25,000 objects in low Earth orbit (LEO). Our radar-powered intelligence protects billions in assets, monitors adversarial behavior, and ensures safe operations for commercial and government missions.

We’re not just building technology, we are redefining global security, safety, and transparency in space. As orbital activity accelerates and threats grow more complex, Leo Labs is a trusted partner for Space Domain Awareness, Space Traffic Management, and Satellite Operations for top-tier space operators and allied defense organizations.

Ifyou'relooking to work on mission-critical challenges at the forefront of aerospace, national security, and AI, your impact starts here.

Role Overview

Leo Labsisseeking a skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will bridge the gap between development and operations, ensuring that our systems are scalable, reliable, and efficient. You will be responsible for automating processes, monitoring system performance, and resolving incidents to enhance our service reliability.

Key Responsibilities
  • System Reliability:
    Design, implement, and maintain scalable and reliable systems.
  • Monitoring and Incident Response:
    Set up monitoring tools and create incident response plans to quickly identify and resolve issues, as well as implementing preventative measures.
  • Automation:
    Develop and maintain scripts and automation tools for deployment, monitoring, and system health checks.
  • Capacity Planning:
    Analyze system capacity and performance metrics to forecast future needs and implement scaling solutions.
  • Collaboration:

    Work closely with development teams to enhance product reliability and streamline the deployment process.
  • Documentation:
    Create and maintain documentation for system architecture, processes, and incident reports.
  • On-Call Support:
    Participate in on-call rotations to provide 24/7 support for critical systems.

Security:
Implement and enforce security best practices across all systems, ensuring compliance with industry standards.

Qualifications
  • Education:

    Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent work experience.
  • Experience:

    5+ years of experience in a Site Reliability Engineering, Dev Ops, or related role.
  • Technical

    Skills:
    • Proficiency in scripting or programming language (e.g., Python, Go)
    • Experience with cloud services (AWS, Azure)
    • Proficiency with containerization (Docker, Kubernetes, ECS)
    • Proficiency in configuration management tools (Terraform, Atlantis, Terragrunt)
    • Familiarity with CI/CD tools (Git Hub Actions, AWS Code Build, CircleCI)
    • Experience with monitoring tools (Grafana, Datadog)
    • Familiarity with database technologies (RDS, Aurora, PostgreSQL)
    • Experience with large-scale distributed systems and microservices architecture.
  • Problem-Solving:
    Strong analytical and problem-solving skills with the ability to troubleshoot complex systems.
  • Communication:
    Excellent verbal and written communication skills, with the ability to collaborate effectively across teams.
  • Ability to obtain a U.S. personnel security clearance.
Preferred qualifications
  • Active TS/SCI clearance
What Success Looks Like

Within 1 month, you’ll:

  • Complete onboarding to understand our business, vision, and team structure.
  • Get familiar with Leo Labs' engineering stack, security posture, and key initiatives.
  • Gain an understanding about how your role fits into Leo Labs broader organization.

Within 3 months, you’ll:

  • Independently deploy infrastructure changes using Infrastructure as Code.
  • Identify key reliability risks and recommend improvements.
  • Improve dashboards, alerts, and operational runbooks.

Within 6 months, you’ll:

  • Optimize infrastructure utilization and cloud costs without compromising reliability.
  • Drive automation that reduces operational toil and improves deployment reliability.

Within 12 months, you’ll:

  • Lead cross-functional initiatives to improve availability, scalability, and operational efficiency.
  • Be a key advisor for site reliability in…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary