More jobs:
Principal Site Reliability Engineer
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-06-02
Listing for:
Upstart
Full Time
position Listed on 2026-06-02
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing
Job Description & How to Apply Below
Requirements
- Bachelor’s degree in Computer Science, Engineering, or Mathematics, or a related field (or its equivalent) + 8 years of experience
- Combined experience with both Software Engineering and Site Reliability Engineering, with a balanced background in both disciplines
- Proven track record as an SRE thought leader and evangelist, driving adoption of reliability best practices across organizations
- Strong communication and mentoring skills to influence engineers across disciplines
- Proficiency in Python, Go, and JavaScript/Type Script
- Proficiency with Infrastructure as Code (Terraform, CDK, Cloud Formation, etc.)
- Experience building internal tooling from scratch in agile development environments
- Expertise with observability, distributed tracing, RUM, LCP, and performance monitoring tools (e.g., Datadog, Prometheus)
- Experience with on‑call and incident management, including large‑scale or ML‑related incidents
- Strong background in automation and building self‑healing systems
- Hands‑on experience with LLM/GenAI to improve SRE efficiency and processes
- Program management skills, including the ability to propose innovative solutions, influence leadership, improve processes, and drive cross‑functional projects to completion
- (Desirable) Experience with service mesh
- (Desirable) Full stack development skills
- (Desirable) Experience building or extending observability platforms
- (Desirable) Background in Development Productivity or Quality Platforms
- (Desirable) Experience in high‑scale SaaS, microservice‑oriented cloud environments
- Upstart’s Site Reliability Engineering (SRE) team owns the reliability, resiliency, and observability of Upstart’s production systems
- We build automation, tooling, and frameworks to ensure our infrastructure is healthy, scalable, and able to support a seamless experience for both engineers and customers
- Our scope includes defining Upstart’s technology operations risk strategy, implementing disaster recovery planning, and setting company‑wide reliability standards
- As a Principal Software Engineer on the SRE team at Upstart, you will serve as a thought leader and SRE evangelist – driving adoption of best practices, mentoring engineers across the organization, and influencing both technical and business decisions
- Your impact will extend beyond SRE into cross‑functional collaboration with Product Engineering, Dev Ex, Development Productivity (Quality), Dev Ops, Data Engineering, and Machine Learning teams to elevate operational excellence across the company
- Lead the definition, advocacy, and adoption of SRE principles across engineering teams
- Partner with leadership to shape long‑term reliability, resiliency, and observability strategies
- Champion distributed tracing, real user monitoring (RUM), and key performance metrics such as Largest Contentful Paint (LCP) to improve system visibility and user experience
- Build and scale self‑healing systems to minimize manual intervention and reduce downtime
- Drive enterprise‑wide improvements to incident response processes, including those related to Machine Learning systems
- Collaborate closely with Development Productivity and Quality teams to improve engineering velocity without sacrificing reliability
- Influence technical and operational roadmaps through data‑driven insights and hands‑on technical contributions
- Own and deliver cross‑functional initiatives from concept through execution, applying program management skills to align stakeholders and achieve results
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×