More jobs:
Site Reliability Engineer - Infrastructure Operations Security Clearance
Job in
San Francisco, San Francisco County, California, 94102, USA
Listed on 2026-08-28
Listing for:
ClearanceJobs Workforce Solutions
Full Time
position Listed on 2026-08-28
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Job Description & How to Apply Below
Clearance Jobs Worforce Solutions is seeking a Site Reliability Engineer – Infrastructure Operations for our client based in San Francisco, CA. Our client is a cutting-edge startup delivering AI-based solutions for critical defense applications serving the U.S. Department of Defense and international defense customers. They operate 24/7 across multiple computing environments—cloud hyperscalers, on-premise GPU clusters, field-deployed systems, and supercomputing centers—with zero tolerance for downtime.
The team is lean and mission-focused, with every member directly impacting the delivery of life-critical systems to their customers. This is a 24/7 on-call infrastructure operations role focused on maintaining mission-critical AI systems at 99% uptime SLA. You'll be the primary responder for infrastructure incidents, monitoring system health, and optimizing existing production pipelines for reliability and performance. This role is NOT about building new infrastructure from scratch.
Our foundations are established. You'll be the operational backbone keeping them running reliably while continually optimizing for performance, cost, and customer SLA commitments.
Key Responsibilities Primary Focus – On-Call Operations & Monitoring (60%)
● 24/7 on-call rotation as the primary responder for infrastructure incidents and production alerts
● Monitor system health and SLA metrics across all customer contracts; escalate to engineering team when needed
● Triage and respond to production incidents: analyze logs/metrics, diagnose root causes, execute remediation, and document post-mortems
● Build and maintain comprehensive observability: whitebox (application-level) and blackbox (system-level) monitoring
● Ensure Datadog and Pager Duty alerting strategies are tuned to catch issues without alert fatigue
● Develop and automate incident response playbooks to minimize MTTR (mean time to recovery)
● Maintain 99% uptime SLA across defense and government contracts
Secondary Focus – Infrastructure Optimization & Reliability (40%)
● Optimize existing infrastructure for performance, cost, and reliability (not redesign from scratch)
● Identify and address infrastructure bottlenecks through capacity planning and performance tuning
● Maintain ETL pipelines and data quality for core forecasting operations
● Collaborate with software engineers on deployment processes and CI/CD improvements
● Document runbooks, playbooks, and operational procedures for team scalability What You'll Bring
Minimum Qualifications
● 5+ years of hands-on experience in SRE, infrastructure operations, or Dev Ops roles
● Python proficiency – it's our primary language; you'll write operational automation tools
● Deep experience with monitoring/observability platforms (Datadog, Prometheus, Grafana, Splunk, ELK, or equivalent)
● Proven on-call incident response experience
● Comfort with 24/7 on-call rotations – you understand production-critical operations and can escalate appropriately
● AWS or Google Cloud infrastructure experience (VPCs, load balancing, databases, serverless)
● Linux systems administration at a production level
● Active Secret clearance or ability to obtain one (required for this role)
Preferred Skills
● Infrastructure as Code (Terraform, Cloud Formation) – nice to have but learnable on the job
● Kubernetes operations (EKS/GKE) – operational expertise more valuable than deep design
● Experience with message queues (SNS/SQS) and caching (Redis)
● GPU resource optimization and ML workload observability
● Experience supporting mission-critical systems for government/defense customers
● Background in incident response and blameless postmortem culture
What We're NOT Looking For
● Someone to redesign infrastructure from scratch (our foundations are solid)
● Pure CI/CD specialists focused on build systems (that's secondary here)
● Infrastructure architects without on-call operations experience
● Candidates uncomfortable with 24/7 on-call responsibilities What You'll Get
● Competitive salary
● Small, world-class engineering team (5 engineers + CTO) – you'll have impact on every decision
● Mission-critical work for U.S. Air Force, Navy, and international defense…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×