Site Reliability Engineer
Listed on 2026-07-14
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Role Overview
Core Scientific is seeking a capable, motivated generalist to work across hybrid cloud and on-premises systems. The role partners closely with application architecture and peer engineering teams, contributing hands‑on across platform engineering, Dev Ops, and SRE.
Responsibilities- Lead end‑to‑end delivery of complex technical initiatives, from problem definition and design through implementation, rollout, and operation.
- Own the design, implementation, and reliability of systems across hybrid cloud and on‑premises environments.
- Take accountability for technical outcomes, including system reliability, scalability, and performance in regulated, change‑controlled environments.
- Drive execution by coordinating work across engineers and teams, delegating effectively while remaining hands‑on where needed.
- Partner with application architecture and peer teams to shape system design and influence technical decisions.
- Build, deploy, and operate infrastructure and applications using automation and infrastructure as code.
- Implement secure, immutable infrastructure using modern tooling (e.g., Terraform, Kubernetes, Helm, Ansible).
- Improve observability, monitoring, and incident response practices.
- Establish and promote best practices for reliability, security, and operational excellence across teams.
- Mentor engineers and contribute to raising the technical bar across the organization.
- Foster open, respectful, and professional communication within the team and with co‑workers, teammates, and leaders across the organization.
- Perform other duties as assigned.
- Bachelor’s degree in Computer Science or a related field, 7+ years of experience, or equivalent demonstrated impact in SRE, Dev Ops, or Infrastructure Engineering.
- Broad technical experience across infrastructure and distributed systems, with the ability to design effective solutions, apply appropriate patterns, and anticipate scaling, reliability, and operational challenges.
- Strong understanding of distributed systems behavior, including application runtime characteristics, service‑to‑service communication, networking, and failure modes in production environments.
- Experience operating in regulated, compliant, or change‑controlled environments.
- Experience working in hybrid environments (AWS preferred; on‑premises infrastructure required).
- Strong experience with Infrastructure as Code, configuration management, and orchestration tools (Terraform, Helm, Kustomize, Ansible).
- Experience with Kubernetes and virtualization technologies.
- Experience with observability platforms (e.g., Datadog), including building monitoring and alerting integrations.
- Experience with build and release systems (e.g., Git Hub Actions, Makefiles, Python tooling).
Site Reliability Engineering Manager
LocationMiami, FL or Austin, TX
TravelOccasional travel may be required as needed.
Work EnvironmentThis job typically operates in a professional office environment and routinely utilizes standard equipment, including laptop computers and smartphones. The role may also travel to data center sites, where the environment may contain loud noise, construction, and other operational elements.
Physical DemandsThe employee is frequently required to sit, stand, walk, use hands, and lift up to 25 pounds.
Position Type / Expected Hours of WorkThis is a full‑time position. General hours and days of work are Monday through Friday, 8:00a.m. to 5:00p.m. The employee is expected to be available generally around U.S. time zones and will be part of an on‑call rotation. The current rotation is 1 week every 5 weeks.
Supervisory ExperienceNo
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).