×
Register Here to Apply for Jobs or Post Jobs. X

Platform Site Reliability Engineer

Job in Lehi, Utah County, Utah, 84043, USA
Listing for: Varite
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 50 - 56.33 USD Hourly USD 50.00 56.33 HOUR
Job Description & How to Apply Below

Platform Site Reliability Engineer

Pay Rate Range: $50 - $56.33/hr hybrid: 2-3 days onsite, open to other office locations - Lehi, UT preferred

Duties:

  • We are seeking a Platform Site Reliability Engineer to ensure the reliability, availability, and operational health of a portfolio of SaaS‐based solutions, including both in‐house (customer‐zero) platforms and vendor‐managed services.
  • This role applies Site Reliability Engineering principles in environments where full‐stack control is not always possible, requiring strong observability practices, operational judgment, and effective collaboration across teams and vendors.
  • You will play a key role in on‐call operations and incident response, establishing meaningful signals from systems you don't fully own and driving practical improvements that reduce risk and customer impact.
  • Success in this role depends on strong communication, creativity in instrumentation and monitoring, and the ability to influence reliability outcomes across organizational and vendor boundaries—leveraging modern, AI‐assisted tooling to improve detection, diagnosis, and learning.

Key Responsibilities:

  • Ensure the reliability, availability, and operational health of a portfolio of SaaS‐based solutions, including vendor‐managed services and in‐house (customer‐zero) platforms.
  • Participate in on‐call rotations and incident response, leading investigation, mitigation, coordination, and post‐incident follow‐up.
  • Establish and maintain effective observability for systems that are not fully owned, identifying practical ways to obtain actionable metrics, logs, and signals from vendor and partner solutions.
  • Use operational data and incident learnings to identify reliability risks and drive targeted improvements that reduce customer impact.
  • Apply appropriate change controls at owned or influenced layers of the stack, balancing reliability, velocity, and business needs.
  • Partner with internal teams and external vendors to communicate expectations, coordinate response and remediation, and influence reliability outcomes.
  • Produce clear incident communications and post‐incident analyses that inform stakeholders and drive lasting improvements.
  • Leverage automation and AI‐assisted tooling to improve detection, triage, and operational efficiency.

Skills:

  • Strong foundation in Site Reliability Engineering practices, including observability, incident response, and reliability measurement.
  • Hands‐on experience operating SaaS or third‐party systems where full‐stack ownership is limited.
  • Deep understanding of monitoring, logging, and alerting, with the ability to design signals that are actionable rather than noisy.
  • Proven incident response experience, including on‐call participation and cross‐team coordination during high‐impact events.
  • Ability to think creatively and pragmatically when instrumenting and improving systems with constrained control.
  • Excellent written and verbal communication skills, especially in high‐pressure incident and vendor‐coordination scenarios.
  • Experience working across organizational and vendor boundaries to resolve complex operational issues.
  • Sound engineering judgment when assessing risk, prioritizing work, and making reliability tradeoffs in production environments.

Education:

Required:

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
  • 3–5 years of experience in Site Reliability Engineering, production operations, or a closely related role.
  • Experience supporting production systems with on‐call responsibilities and incident response expectations.
  • Strong experience working with observability data (metrics, logs, alerts) to diagnose issues and drive improvements.
  • Comfort using automation and AI‐assisted tools as part of everyday operational workflows.

Preferred:

  • Experience supporting enterprise‐scale SaaS platforms or shared services.
  • Prior experience working directly with vendors to resolve reliability or operational issues.
  • Familiarity with cloud‐based and distributed system architectures.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary