Sr. Principal Engineer — Platform
Listed on 2026-09-12
-
Software Development
Cloud Engineer - Software, Software Architect, DevOps
We are looking for a Sr. Principal Engineer to provide technical leadership and make significant contributions to the architecture and development of Rafay's multi-tenant cloud and AI infrastructure platform.
Rafay operates at the intersection of distributed systems, Kubernetes, virtualization, infrastructure automation, observability, and security. This role provides an opportunity to build foundational technologies used to operate complex cloud and AI infrastructure environments at scale.
As a Sr. Principal Engineer, you will serve as a technical leader and hands-on architect
, setting technical direction for critical platform capabilities and multiplying the effectiveness of engineering teams across the organization.
You will be expected to move comfortably between architecture and implementation, reason about complex distributed systems and failure modes, and design platforms that are scalable, highly available, observable, secure, and suitable for enterprise and regulated environments.
Responsibilities- Design and implement core architectural components for critical services within a large-scale, multi-tenant distributed platform.
- Set the technical vision and long-term architectural direction for major platform areas and drive alignment across engineering teams.
- Design highly modular, scalable, resilient, secure, and maintainable distributed services.
- Architect platform capabilities spanning cloud infrastructure, Kubernetes, virtualization, observability, security, and automation
. - Drive the design of modern observability capabilities covering metrics, logs, traces, events, infrastructure telemetry, and service health.
- Develop approaches for correlating information across multiple infrastructure and application layers to improve troubleshooting and operational reliability.
- Design service health monitoring, alerting, SLI/SLO frameworks, synthetic monitoring, and proactive validation capabilities.
- Develop capabilities that improve incident detection, troubleshooting, root-cause analysis, and operational automation
. - Help establish architectures for safe, controlled, and auditable automation of infrastructure operations.
- Define and drive platform security architecture, including identity, authorization, secrets management, workload isolation, network security, secure APIs, and privileged operations.
- Design systems suitable for high-security and regulated environments
, including government and enterprise deployments. - Partner with security and compliance teams to translate compliance requirements into scalable platform capabilities.
- Contribute to architectures involving confidential computing, trusted execution environments, hardware-backed security, attestation, and protection of sensitive workloads and data
. - Participate in security architecture reviews and threat modeling for distributed and multi-tenant systems.
- Ensure systems provide strong auditability for administrative actions, configuration changes, and automated operations.
- Perform R&D and feasibility analysis on emerging technologies related to cloud infrastructure, AI infrastructure, observability, distributed systems, and security.
- Assist engineering and operations teams with diagnosing complex production issues and drive systemic improvements based on lessons learned.
- Lead architecture reviews, design reviews, and code reviews and help establish engineering standards across teams.
- Mentor engineers and raise the technical bar across the engineering organization.
- Collaborate with engineering, product management, security, SRE, QA, and customer-facing organizations.
- Champion a customer-focused engineering culture where production experience informs improvements in reliability, usability, security, and automation.
- 12+ years of experience designing, building, and delivering large-scale enterprise software platforms.
- Demonstrated experience operating as a senior technical architect or Principal-level engineer responsible for significant platform architecture decisions.
- Deep understanding of distributed systems fundamentals, including:
- Scalability
- High availability
- Resiliency
- Distributed state
- Concurrency
- Failure handling
- Performance
- Expert knowledge of one or more programming languages, preferably:
- Golang
- Python
- Strong experience designing and developing microservices and distributed control-plane systems.
- Excellent troubleshooting and debugging skills across complex production environments.
- Hands-on experience building services on public cloud platforms such as AWS, Azure, or…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).