Site Reliability Engineer - FedRAMP
Listed on 2026-07-27
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for aSite Reliability Engineer - FedRAMP based in United States.
This role offers the opportunity to strengthen the reliability and operational excellence of a large-scale SaaS platform supporting government and sovereign cloud environments.
You will work at the intersection of cloud infrastructure, security, automation, and software engineering to improve system resilience.
The position focuses on building reliable services, enhancing observability, and supporting incident response in compliance-driven environments.
You will collaborate with experienced engineers across engineering, security, and operations teams to solve complex reliability challenges.
The role provides hands‑on ownership of infrastructure improvements while contributing to long‑term reliability strategies.
This is an ideal opportunity for an engineer who enjoys solving ambiguous problems, improving systems, and working in a high‑impact cloud environment.
The Site Reliability Engineer will help maintain and improve the operational foundation of a secure, cloud-based platform. This role requires strong technical execution, proactive problem-solving, and close collaboration across engineering and operational teams.
- Gain deep understanding of platform workloads, dependencies, and operational workflows through documentation, code analysis, and collaboration with subject matter experts.
- Create and maintain operational documentation, including runbooks, incident guides, onboarding resources, and knowledge-sharing materials.
- Participate in incident response activities, including investigation, mitigation, root cause analysis, and post‑incident improvements.
- Support the implementation and maintenance of reliability practices, including SLIs, SLOs, error budgets, and availability improvements.
- Improve system observability by developing monitoring, alerting, dashboards, and instrumentation strategies.
- Reduce operational complexity through automation, tooling improvements, and elimination of repetitive tasks.
- Support infrastructure delivery through infrastructure-as-code, CI/CD pipelines, deployment workflows, and configuration management.
- Contribute to secure and compliant infrastructure changes within regulated environments.
- Collaborate with engineering, security, compliance, and operations teams to improve reliability, communicate risks, and resolve technical challenges.
- Participate in on‑call rotations and help ensure platform stability and resilience.
The ideal candidate brings strong software engineering fundamentals, cloud infrastructure experience, and the ability to operate effectively in regulated environments. They are comfortable investigating complex systems, improving reliability practices, and collaborating across technical teams.
- 3+ years of experience in software engineering, including at least 1 year working in Site Reliability Engineering, Platform Engineering, or Dev Ops roles supporting cloud-hosted services.
- Experience with cloud infrastructure platforms such as Azure or similar cloud providers.
- Familiarity with compliance-focused environments such as government, FedRAMP, CMMC, financial services, or healthcare industries.
- Ability to understand and troubleshoot application code to investigate system behavior independently.
- Experience with observability and monitoring tools such as Prometheus, Grafana, Open Telemetry, or ELK stack.
- Hands‑on experience with infrastructure-as-code tools such as Terraform, Terragrunt, or Pulumi.
- Experience with container orchestration platforms, particularly Kubernetes.
- Experience managing CI/CD workflows using tools such as Git Hub Actions, Azure Dev Ops, Git Lab CI, or ArgoCD.
- Strong programming skills in languages such as Type Script, JavaScript, Go, Java, C#, or similar.
- Understanding of distributed systems concepts, networking fundamentals, and cloud reliability principles.
- Strong written and verbal communication skills with the ability to explain technical concepts clearly.
- Experience with government or sovereign cloud environments, SaaS platforms,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).