Senior Platform Engineer – AI/ML Infrastructure & Reliability
Listed on 2026-08-22
-
Software Development
$210,000–$260,000 Base + Equity
About the CompanyOur client is a fast-growing San Francisco technology company building an advanced AI platform on high-performance distributed infrastructure.
The engineering organization is small and highly technical. Engineers work closely with software teams, applied scientists, product, and leadership, with significant individual ownership and a short path from identifying a problem to deploying a production solution.
About the RoleWe're looking for a senior platform engineer who is fundamentally a builder
.
This is not a production operations role
. We're looking for someone who has personally owned substantial technical projects end-to-end:
You'll build foundational platform and infrastructure capabilities supporting advanced AI, data, and distributed workloads. Some projects will start with well-defined requirements; others will begin with an ambiguous technical problem that you will help define and solve.
The strongest candidates have experience in fast-growing startup environments
, where engineers operate with significant autonomy, responsibilities are broad, and building from scratch is part of the job.
- Own platform and infrastructure projects from initial problem through production
- Design and build Kubernetes and container-based infrastructure
- Develop internal tooling, automation, and platform software
- Build reusable Infrastructure-as-Code and deployment systems
- Create CI/CD, Git Ops, and developer self-service capabilities
- Solve technical problems across Linux, networking, storage, cloud, and distributed systems
- Build reliability and observability into systems from the beginning
- Partner directly with engineers, scientists, product teams, and technical leadership
- Own and evolve the systems you build after launch
- 8+ years in platform engineering, infrastructure engineering, production engineering, SRE, Dev Ops, systems engineering, or related software engineering
- Evidence of building or substantially redesigning systems—not simply operating them
- Strong Linux and systems fundamentals
- Terraform or comparable Infrastructure-as-Code expertise
- Python, Go, or comparable programming/automation experience
- Experience building CI/CD, deployment automation, or developer-platform capabilities
- Strong troubleshooting and distributed-systems fundamentals
- Ability to make architectural tradeoffs across performance, reliability, security, cost, and complexity
- Comfort operating independently when requirements are incomplete or changing
- Fast-growing startup experience
- Founding or early infrastructure/platform engineering experience
- GPU, NVIDIA, or HPC infrastructure
- AI/ML training or inference infrastructure
- Distributed compute or high-throughput systems
- Kafka, Spark, Airflow, Ray, or Databricks
- Advanced networking or distributed storage
- Git Ops, ArgoCD, Helm, or service mesh technologies
You enjoy building more than maintaining
. You want meaningful technical ownership, are comfortable moving outside a narrow specialty, and don't need someone assigning your next ticket.
You can take an ambiguous problem, determine what needs to be built, evaluate the tradeoffs, design the solution, implement it, and take responsibility for the result.
Why Join- Build foundational systems rather than simply maintain existing infrastructure
- Own technically meaningful projects from concept through production
- Work on challenging AI, distributed-systems, data, and compute problems
- Join a small engineering organization where individual contributions matter
- Work directly with highly technical engineers, scientists, and leadership
- $210,000–$260,000 base salary plus equity
Work Authorization: Candidates must be currently authorized to work in the United States. This position is not eligible for new or future employer-sponsored work authorization.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).