Principal Engineer, Distributed Systems
Listed on 2026-09-20
-
IT/Tech
Core Weave is The Essential Cloud for AI™. Built for pioneers by pioneers, Core Weave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, Core Weave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, Core Weave became a publicly traded company (Nasdaq: CRWV) in March 2025.
Learn more at
We are looking for a Principal Engineer to provide technical leadership across Security Products. This is a senior individual-contributor role for an engineer who can define architecture, guide execution across multiple teams, and solve complex distributed-systems problems in security-critical infrastructure. The systems you help design will support demanding production requirements: four nines of availability, high scalability, consistently low latency, strong security boundaries, and safe behavior under partial failure in a multi-region setup.
You will work across the full lifecycle of these systems, from architecture and technical strategy through implementation guidance, operational readiness, incident learning, and long-term evolution. Your work will span across high-scale authorization systems, anomaly detection systems, bot-defense and security token service (STS), API authentication gateway, and the shared infrastructure required to operate these capabilities reliably across regions and deployment environments.
You will partner with engineering, security, infrastructure, networking, platform, product, and customer-facing teams. Success in this role requires both deep technical judgment and the ability to create alignment, raise engineering standards, and make complex architecture understandable and actionable for others.
- Establish architectural approaches for distributed systems - multi-region operation, including service placement, failover, replication, traffic management, disaster recovery, and regional independence.
- Drive decisions around consistency models, caching, invalidation, propagation, revocation, idempotency, concurrency, and ordering where correctness and security are critical.
- Design for data residency, tenant isolation, trust boundaries, blast‑radius reduction, and controlled handling of sensitive security data.
- Define fault‑tolerance strategies for dependencies, networks, regions, storage systems, and control‑plane components, including graceful degradation and safe recovery.
- Design systems that meet four‑nines availability goals while maintaining predictable low latency and high throughput under normal operation, traffic spikes, and partial failures.
- Improve the reliability and operability of critical services through SLOs, error budgets, metrics, logs, traces, audit events, alerting, incident response, and post‑incident learning.
- Guide teams through architecture reviews, design reviews, implementation tradeoffs, capacity planning, load testing, performance analysis, and production readiness assessments.
- Mentor senior and staff engineers, develop technical talent, and raise the quality of engineering practice across the organization.
- Communicate architecture, tradeoffs, risks, and recommendations clearly to technical and executive audiences.
- Lead the design of high-scale authorization systems that support policy authoring, policy evaluation, access‑control enforcement, auditability, and integration across many services and tenants.
- Provide technical direction for a Security Token Service, including token issuance, validation, lifecycle management, trust relationships, key rotation, revocation, and secure service‑to‑service access.
- G…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).