×
Register Here to Apply for Jobs or Post Jobs. X

Principal Site Reliability Engineer

Job in Hillsboro, Washington County, Oregon, 97104, USA
Listing for: Saviynt
Full Time position
Listed on 2026-09-22
Job specializations:
  • Software Development
    Cloud Engineer - Software, Backend Developer, DevOps
Salary/Wage Range or Industry Benchmark: 260000 - 275000 USD Yearly USD 260000.00 275000.00 YEAR
Job Description & How to Apply Below

Why Join Saviynt

  • Work on a mission-critical SaaS platform used by global enterprises
  • Solve complex reliability challenges at scale
  • Influence architecture and engineering culture at a company level
  • Competitive compensation, benefits, and growth opportunities
Security & Compliance

This role requires compliance with Saviynt’s information security and privacy policies, including annual security training

What You Will Be Doing
  • In this pivotal role, you will be instrumental in designing, building, and maintaining the shared infrastructure services and platforms that our product and application teams will depend on
  • You will focus on creating reusable, reliable, and scalable solutions that abstract away complexity, enabling other teams to focus on their core business logic and deliver features faster in a multi-cloud environment
  • Design and build core platform components and shared infrastructure services that other development teams will integrate with and leverage to deploy and operate their applications
  • Architect, implement, and manage highly available and scalable Kubernetes platforms as a service for internal consumers
  • Develop robust, internal-facing tools and automation for infrastructure provisioning and management primarily using Go (Golang)
  • Architect and optimize foundational solutions within Cloud environments (AWS, Azure, etc.), focusing on creating reusable patterns and modules for other teams
  • Design and implement shared Event-Driven Architecture components and messaging platforms using technologies like Kafka or Google Pub/Sub that product teams can easily utilize
  • Develop and maintain robust CI/CD pipelines (e.g., Git Lab CI and ArgoCD) as a service, providing standardized and automated deployment workflows for various development teams
  • Design and build resilient Distributed Systems components that serve as building blocks for other applications, focusing on reliability, fault tolerance, and performance
  • Manage and optimize our shared infrastructure across Multi-Region Cloud Environments, ensuring that platform services are globally available and performant for all consumers
  • Establish and enhance centralized Observability and Monitoring platforms and tools that provide self-service insights for consuming teams
  • Define and implement clear, well-documented RESTful API designs for the infrastructure services you build, ensuring ease of integration for internal clients
  • Implement and manage Service Mesh (e.g., Envoy, Istio) capabilities, providing traffic management, security, and policy enforcement as a shared platform for services
  • Design, implement, and optimize highly available Relational Database services or shared data platforms for broad organizational use
  • Collaborate closely with product development teams to understand their infrastructure needs and pain points, providing technical guidance and support
  • Participate in on-call rotations to support the critical shared infrastructure you build
What You Bring
  • 1+ years of experience as a Principal SRE with a strong focus on building tools and services for other engineers
  • Deep expertise with Kubernetes in production environments, particularly in providing it as a platform(i.e single tenant and multi-tenant deployment architectures)
  • Strong programming skills in Go (Golang) and Python, with experience building robust, maintainable backend services and automation
  • Extensive hands-on experience with at least one major Cloud Provider (AWS, GCP, or Azure); multi-cloud experience is a strong plus, especially in building abstractions over them
  • Proven experience designing and implementing Event-Driven Architecture and message queuing systems (e.g., Kafka, RMQ, NATS) as shared services
  • Solid understanding and practical experience with CI/CD pipeline tools (especially Git Lab CI) and experience establishing automated delivery processes for other teams
  • Demonstrable experience designing and operating Distributed Systems, with an understanding of patterns for creating reliable, shared components
  • Familiarity with Multi-Region Cloud Environments and strategies for building globally distributed and highly available platform
  • Proficiency in establishing and utilizing comprehensive Observability and Monitoring platforms (e.g., Prometheus, Grafana, ELK stack, Datadog) for shared infrastructure
  • Strong experience with RESTful API design principles and building well-documented, consumable APIs
  • Knowledge of Service Mesh concepts and practical experience with solutions like Istio in a platform context
  • Hands-on experience with…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary