Lead Infrastructure & Platform Engineer
Listed on 2026-08-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Infrastructure
Studyfetch Lead Platform Engineer
Study Fetch is the #1 AI-native learning platform globally, transforming how millions of students learn through personalized AI-powered education. We're growing fast with backing from top-tier investors and a mission that's redefining the future of education and ethical learning.
We're a technology company building AI-native learning products used by more than seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both run on the Learn Engine, the intelligence that moves a learner from initial understanding to demonstrated mastery. We work with partners like NVIDIA to bring responsible, learning-first AI to the students who need it most.
All of that runs on a mature, infrastructure-as-code platform: cloud, Kubernetes, networking, deploy tooling, databases, and the GPUs serving that powers our AI. We're hiring a lead to own that platform and take it further as we scale.
This is a high-ownership role. You'll be the person the rest of engineering depends on to ship, and you'll set the patterns for how we deploy, secure, and operate. You'll also lead the next chapter: we're building out our own GPU infrastructure inside a colocation facility, which means real networking, capacity planning, and bare-metal platform work sitting alongside our cloud footprint.
If you want a surface where the decisions are yours and the impact reaches millions of learners, this is it.
Every learner deserves the chance to succeed. Study Fetch started with one idea: high-quality, personalized learning should be within reach for anyone, at any stage of life. Honen carries that belief into the workforce.
- Accessible to everyone. Learning should reach every person, whatever their background, role, or resources.
- Meet people where they are. Every course adapts to a person's pace, their level, and the way they learn best.
- Learning never stops. From a first job to a new career, people keep growing at every stage of life.
We hire people who share this conviction. The work is demanding and the hours can be long, and what sustains you through it is caring whether a real student finally understands the material.
You'll own the platform, end to end. Our infrastructure-as-code (Pulumi/Type Script on GCP), the Kubernetes clusters, the Shared VPC and networking, the CI/CD and deploy tooling, secrets, and the databases behind them. You own how the whole thing fits together and how the team ships on top of it.
You'll be responsible for reliability and the on-call that follows. Monitoring, alerting, incident response, and the postmortems that make the next incident less likely. When production has a bad night, you're the person who understands why and makes sure it doesn't repeat.
You'll help design and build our own GPU infrastructure: hardware and capacity planning, physical and virtual networking, the platform layer that makes those GPUs usable for model serving, and the path that connects it cleanly to our cloud environment. This is a meaningful part of the role, and it's new ground for the company.
You'll manage the GPU serving platform. The clusters and pipelines that run our self-hosted models and speech/ASR workloads, keeping latency, cost, and utilization where they need to be for real users at scale.
You'll maintain the security and compliance posture. We hold ourselves to a real bar (SOC 2, continuous scanning, least-privilege IAM). You keep the platform audit-ready without slowing the team down.
The patterns you set for how we deploy, secure, and operate are the ones everyone else follows. As the work grows, you'll shape and help grow the team that does it with you.
You're a strong fit if either of these is true:
- 7+ years building and operating production infrastructure, with real depth in cloud platform, Kubernetes, and networking, or
- You were the founding or early infrastructure engineer who owned a significant share of a real platform. Fewer years on paper, but you took production infrastructure from early and messy to reliable and are able to demonstrate it.
Beyond that:
- You've run production infrastructure that real users depend on, not lab setups. You can walk us through a platform you built or operated, which parts were yours, an incident that went badly, and what you changed because of it.
- You have hands-on experience with physical or colocation infrastructure. Bare-metal provisioning, datacenter or colo networking, hardware and capacity planning, GPU fleets, or standing up a hybrid of on-prem and cloud. This is the newest part of the role, and experience here is a real differentiator.
- You're fluent in modern cloud and Kubernetes. Infrastructure-as-code is how you work, not a thing you tolerate. You have opinions about ownership boundaries, blast radius, and what belongs where.
- You use AI every day and have informed opinions about it. You understand the demands of serving models in production, and genuine curiosity is the one thing we can't teach.
- You've worked through launch…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).