×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer Remote ​/ Telecommute Jobs

Remote / Online - Candidates ideally in
Washington, District of Columbia, 20001, USA
Listing for: Clearance Jobs
Full Time, Remote/Work from Home position
Listed on 2026-08-16
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
Job Description & How to Apply Below

Site Reliability Engineer

On behalf of our client, Clearance Jobs Workforce Solutions is seeking a Site Reliability Engineer. This full-time (direct hire) position is located in San Francisco, CA. Work is performed 100% on-site. Our client is willing to consider candidates located near Omaha, NE, Boston, MA, and in Washington, D.C. metro area with expectation of quarterly travel to CA and NE. Candidates must be able to work on our client's W2 (annual salary + benefits package).

Due to government contract requirements, United States Citizenship and an active DOD Secret security clearance is required. No 3rd party Corp to Corp inquiries will be considered.

Summary:

Our client is a cutting-edge startup focused on delivering AI-based weather forecasting solutions to enhance climate resilience and public safety. To deliver advanced weather forecasts, they use numerous cutting-edge computing environments, including cloud-based hyperscalers, on-premise GPU clusters, field-deployed computers, and large supercomputing centers. They work closely with the Department of Defense (DoD) and must adhere to strict compliance and security standards.

The team thrives in a dynamic, fast-paced environment, and every member plays a critical role in driving our mission forward.

Responsibilities:

  • Scaling Production Environment:
    • Design and implement scalable infrastructure solutions to support growing business needs.
    • Optimize system performance and availability through capacity planning and performance tuning.
  • Monitoring Stack Improvement:
    • Develop and maintain whitebox (application-level) and blackbox (system-level) monitoring systems.
    • Ensure comprehensive observability through the integration of metrics, logging, and tracing.
    • Utilize tools such as Datadog and Pager Duty to establish reliable alerting and incident response processes.
  • ETL Pipeline Management:
    • Design, launch, and maintain robust ETL pipelines to support data-driven operations.
    • Collaborate with data teams to ensure data quality and pipeline reliability.
  • CI/CD and Build Systems:
    • Implement and manage continuous integration and continuous deployment pipelines.
    • Improve developer productivity by maintaining reliable build systems and workflows.
  • On-call Responsibilities:
    • Participate in on-call rotations to ensure high availability and timely incident resolution.
    • Develop and automate incident response playbooks to minimize downtime.

Required:

  • 6+ years of experience in SRE or Dev Ops roles.
  • Proficiency with infrastructure as code tools, particularly Terraform.
  • Experience with AWS &/or Google Cloud.
  • Strong background in software engineering and CI/CD pipeline management.
  • Experience with on-call operations and incident management.
  • Familiarity with monitoring and alerting tools such as Datadog and Pager Duty.
  • Knowledge of cloud-native services, including SNS/SQS and Redis.
  • Experience with ML experiment tracking and GPU optimization is a plus.

Desired:

  • Expertise in managing and scaling distributed systems.
  • Strong understanding of networking, security, and Linux systems.
  • Experience in automating infrastructure and deployment processes.
  • Familiarity with message queues (SNS/SQS) and caching systems (Redis).
  • Knowledge of ML workflows and GPU resource management.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary