Site Reliability Engineer — AI GPU Infrastructure
Listed on 2026-09-25
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, Network Engineer
United States Digital Space LLC is seeking a Site Reliability Engineer (STARSHIELD) to design, operate, and scale GPU/CPU infrastructure for Top Secret datacenters and AI clusters. The role focuses on automation, on-prem Kubernetes, and collaboration with AI teams to deliver scalable, reliable software products.
The position requires SRE/Dev Ops experience, Linux proficiency, and familiarity with Terraform/Ansible. A TS/SCI-like clearance and willingness to travel are important for success.
This posting is for the Site Reliability Engineer — AI GPU Infrastructure role at United States Digital Space LLC, based in United States.
Step into the Site Reliability Engineer — AI GPU Infrastructure role at United States Digital Space LLC in United States and grow with us.
Please review the full job details above before applying.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).