ASKUSRSite Reliability Engineer; SRE
Listed on 2026-09-06
-
IT/Tech
Systems Administrator, SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Location: Northern
Work Location:
Onsite - California
Schedule:
Full-Time | 5 Days Per Week | Midnight-8:00 AM (Owl Shift)
This position does not offer sponsorship. Must be authorized to work in the United States.
Position OverviewEssnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance computing (HPC) and data environment supporting scientific research for the U.S. Department of Energy (DOE) Office of Science.
The Site Reliability Engineer will work as part of a 24/7 operations environment responsible for maintaining the accessibility, reliability, security, and operational health of large-scale computing and data systems.
This is a highly hands-on position combining Linux systems administration, infrastructure monitoring, incident response, programming and scripting, automation, networking, Service Now, and physical data center operations.
Key Responsibilities- Monitor high-performance computing systems, storage infrastructure, networks, and other data center and facility-related systems.
- Review and respond to infrastructure and system alerts, perform initial triage, and engage appropriate on-call personnel when escalation is required.
- Respond to alerts across multiple systems to help ensure monitoring and data collection remain operational 24/7.
- Troubleshoot system, application, network, monitoring, and infrastructure issues affecting system reliability.
- Develop solutions that improve operational processes, prevent recurring issues, and automate responses to routine service conditions.
- Identify opportunities to improve monitoring capabilities, alerting, incident triage, and operational automation.
- Develop and maintain tools within the monitoring pipeline in collaboration with operations personnel.
- Develop software and integrations capable of generating alerts and notifications from HPC system APIs into monitoring pipelines.
- Build and maintain application and tool configurations to ensure reliable operation as data volumes and user demands increase.
- Utilize Service Now to support incident management, trouble-ticketing, operational workflows, and service management activities.
- Collaborate across technical teams to identify and resolve operational bottlenecks and maintain system reliability.
- Coordinate with technical groups during center-wide maintenance activities.
- Manage diagnostic, monitoring, and notification software during planned maintenance periods.
- Perform regular physical and logical walkthroughs of the data center floor.
- Monitor environmental conditions, power distribution units (PDUs), cooling infrastructure, and other facility systems supporting reliable data center operations.
- Maintain accurate trouble-ticket documentation for outages, incidents, maintenance activities, troubleshooting actions, and operational updates.
- Analyze problems of varying complexity and evaluate technical data to determine appropriate troubleshooting and remediation methods.
- Exercise independent technical judgment when selecting methods and approaches for resolving operational issues.
$80.00 per hour
The anticipated pay rate for this position is $80.00 per hour. Actual compensation may be determined based on job-related factors including experience, qualifications, skills, contractual requirements, and applicable law.
Equal Employment OpportunityEssnova Solutions, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, creed, sex, pregnancy, childbirth or related medical conditions, sexual orientation, gender, gender identity or expression, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, marital status, military or veteran status, or any other characteristic protected by applicable federal, state, or local law.
Essnova Solutions, Inc. is committed to providing reasonable accommodations to qualified individuals with disabilities and applicants with disabilities throughout the recruitment and employment process.
Required Qualifications- 5+ years of relevant professional experience in Site Reliability Engineering, systems/infrastructure engineering, Dev Ops, data center operations, HPC operations, network/system operations, or a closely related technical environment.
- Strong hands-on experience working with Linux, including Linux shell and command-line environments such as SSH.
- Programming and/or scripting experience using one or more languages such as:
- Python
- C
- C++
- Perl
- Java
- Comparable scripting or programming languages
- Knowledge of standard software development practices.
- Experience supporting large-scale IT infrastructure, highly available systems, data centers, critical installations, or comparable technical environments.
- Knowledge of large data communications networks and common network protocols.
- Network security experience, including knowledge of firewalls and access control lists…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).