×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer III

Job in Newport News, Virginia, 23600, USA
Listing for: Phase2 Technology
Full Time position
Listed on 2026-09-13
Job specializations:
  • IT/Tech
    Systems Engineer, SRE/Site Reliability, IT Project Manager, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 118400 - 186850 USD Yearly USD 118400.00 186850.00 YEAR
Job Description & How to Apply Below

At Jefferson Lab,you'llchampioncutting-edge science and operational excellence while shaping the future of discovery. Join us and make your mark - where excellence meets purpose, andgreat mindstruly
matter.

The good-faith pay range for this role is $118,400 - $186,850 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.

What your job will be like:

As Lead Site Reliability Engineer on the High Performance Data Facility (HPDF) team, you will play a critical role in establishing and running the reliability practice for the facility's first systems on its path to operations. This is a technical role with manager responsibilities: you will supervise and develop a small team of site reliability engineers, and you will also design and build systems yourself, hands on in the code, the monitoring stack, and the incident response.

You will design how the facility stays available and recovers, define and report on the service level objectives that measure how well it serves its users, serve as incident commander for significant incidents, and work day to day with staff at both Jefferson Lab and Berkeley Lab. HPDF is still in design, so there is room in this role to grow into influencing the technology choices the facility is built on.

The users you support are research physicists and computational scientists, and helping them succeed is a core measure of this role.

In this job you will:
  • Lead the design, implementation, and operation of monitoring, logging, alerting, and diagnostic tooling for HPDF compute, storage, network, and facility systems, contributing directly to that work as well as directing it.
  • Supervise, mentor, and develop a small team of site reliability engineers: assign and review work, set expectations, give regular feedback, support technical growth, and plan and estimate the multi-person efforts assigned to the team.
  • Establish and maintain the facility's operational framework, including on-call and escalation structure, incident management, change management, and scheduled maintenance, and keep operational records, runbooks, and documentation current.
  • Design the facility's resilience model, including failure domain isolation, redundancy, graceful degradation, and disaster recovery objectives for a geographically distributed facility, and validate that design through testing.
  • Define, implement, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in collaboration with the architecture team and scientific stakeholders, and hold facility operations to them.
  • Serve as incident commander for significant incidents, own the postmortem process, and drive root cause prevention back into the design and operation of the systems.
  • Drive reliability improvement through automation, process optimization, and the elimination of manual operations, using Python, Go, or shell and standard software development practices.
  • Partner with the architecture team on HPDF technology selection from a reliability standpoint, lead evaluations of vendor and open source technologies, and represent HPDF site reliability engineering in the Berkeley Lab partnership.
Additional Responsibilities
  • Participate in an on-call rotation as the facility moves toward operations.
Lead
- Supervisory
- Management
  • Supervises a team of site reliability engineers
  • Assigns and reviews work, sets performance expectations, provides regular feedback, conducts performance discussions, and supports the technical development of the team.
  • Participates in hiring for the group. Does not hold fiscal or budget authority.
Experience
  • Required:

    10 or more years experience in Site Reliability Engineering, Dev Ops,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary