More jobs:
Sr. Site Reliability Engineer, AI Infrastructure; Starshield
Job in
Palo Alto, Santa Clara County, California, 94302, USA
Listed on 2026-10-08
Listing for:
SpaceX
Full Time
position Listed on 2026-10-08
Job specializations:
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Unix/Linux
Job Description & How to Apply Below
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.
SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation. Starshield is the world’s largest US government satellite constellation and is tasked with providing immediate access to critical intelligence and national security data for the US government anywhere on the globe. We design, build, test, and operate all parts of the system – receivers that allow users to connect within minutes, and the software that brings it all together.
We’ve only begun to scratch the surface of Starshield's global impact and are looking for best-in-class engineers to help us further our ambitious goals.
As an engineer focused on Starshield's software and GPU infrastructure, you will design, operate and scale the infrastructure which supports critical national security missions. These positions cover a variety of areas ranging from Site Reliability Engineering, Developer Operations, and GPU platforms. You will develop automation to deploy and manage on-premise compute resources, create highly scalable and maintainable software products, and directly collaborate with engineering across the board.
RESPONSIBILITIES:
Manage GPU/CPU infrastructure deployments to Top Secret data centers
Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
Develop automation to deploy and manage on-premise KubernetesAI clusters, and operating systems
Deploy and manage core infrastructure such as databases, monitoring and distributed storage
Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement
Monitoring and alerting supporting systems to have high availability
Identify areas for improvement and create innovative solutions that enable high system availability
Mentor and train junior engineers
As a senior engineer you must lead the team to technical excellence – your decisions guide the team
BASIC QUALIFICATIONS:
Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 5+ years of professional experience with Linux operating systems; OR 7+ years of professional experience in software, Dev Ops, or site reliability engineering in lieu of a degree5+ year of experience with Kubernetes5+ year of experience managing Linux operating systems
Experience with Terraform, Ansible, or other infrastructure tools
Experience with containerization technologies (i.e. OCI containers, Kubernetes)
Experience scripting in Bash, Python, or other similar languages
Development experience in Python, C++, or GoPREFERRED
SKILLS AND EXPERIENCE:
5+ years of experience with Python and Python-based development frameworks
Experience managing Kubernetes clusters, not just using them Knowledge of Linux boot process and systems configuration
Deep understanding of testing, continuous integration, build, deployment & continuous monitoring
Understanding of relevant build technologies, such as Bazel and Makefiles Focus on performance bottlenecks and performance improvement techniques
Understanding of distributed databases and data modeling
Experience with automatically managing thousands of servers (eg: Terraform or Ansible)
Strong networking knowledge of…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×