×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Pacifica, San Mateo County, California, 94045, USA
Listing for: Pacificacontinental
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

Our engineering team has built the largest private Medicare marketplace in the country. We passionately focus on the continuous improvement of the systems we build.

We have spent many years growing and fostering a Dev Ops culture by bridging the divide between our Software and Infrastructure Engineering departments. We want the cross-functional teams that we are building to include Site Reliability Engineers. We operate in a complex, multi-tenant, hybrid cloud and on-premises infrastructure that spans both the Windows and Linux OS. We strive for security, reliability, and automation in line with Dev Ops and Site Reliability Engineering principles.

If you are passionate about learning and improvement through metrics and automation, and passionate about engendering that mindset in others, we want to hear from you.

About the role:

Maintains shared cloud resources in use by numerous software engineering teams within our business unit. We aim to enable software engineering teams to build cloud native applications that adhere to security and regulatory requirements with limited hand holding by our cloud engineers. We do still have a fair number of applications hosted in on-premise data centers, which we aim to support migrating to the cloud.

Requirements:

Hands-on Engineering

5+ years of hands-on experience with a majority of the following technologies, along with a willingness to become proficient in the remaining areas:

  • Windows and Linux Servers
  • VMware
  • Cloud platforms, preferably with Azure
  • Active Directory
  • Secrets management with Consul and Vault or similar systems
  • Configuration management tools like Salt, Ansible and Terraform
  • Firewalls and load balancers such as F5
  • Web servers, including IIS andNGINX
  • Database Server Infrastructure like Microsoft SQL Server and PostgreSQL
  • Application Performance Monitoring with tools like New Relic
  • Infrastructure monitoring with tools like Sensu, Solar Winds, Nagios, or Azure App Insights
  • CI/CD tools like Team City, Octopus Deploy, Concourse, Azure Dev Ops, or Git Hub Actions
  • Log Aggregation tools like Sumo Logic or Splunk
  • Network theory and protocols such as DNS, DHCP, proxy servers, and firewalls
  • Security operations with tools for SAST, DAST, RAST, and WAF
  • Infrastructure as Code or automation experience.

Proficiency, high-comfort, and familiarity with:

  • One or more scripting languages, such as Power Shell and BASH
  • Command line tools such as (git, netcat, npm, terraform, etc.)
Responsibilities
  • Make improvements to internal processes to reduce lead time and increase deployment frequency
  • Identify improvements to the quality, security, and performance of our infrastructure
  • Increase the velocity with which teams deliver, leveraging expertise from various functional disciplines
  • Identify how to remediate production incidents more quickly and safely while reducing the frequency of outages
  • Actively engage with other teams and departments to collaborate on best practices and implementation strategy
  • Adhere to and advocate for best practices, including Infrastructure as Code, monitoring, high availability,disaster recovery,security, and Dev Ops methodologies
  • Create SLIs, SLOs, and SLAs
  • Contribute to capacity planning, advise and consult with teams who will be load/stress testing
  • Keep up with industry innovations, recommending new tools or practices when appropriate
  • Actively mentor peers, developing their expertise and inspiring others to innovate
  • Provide timely assistance and remediation solutions during critical situations and production incident
  • Document and share “lessons learned” from production, including root cause analysis
  • Explore new ways of improving communication between other Site Reliability Engineers and with other teams
  • Write and maintain architectural, stakeholder, and policy documentation

Attach resume as .pdf, .doc, .docx, .odt, .txt, or .rtf (limit 5MB) orpaste resume

References:
Please enter names and contact information:*

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary