AI Product Engineer Houston San Francisco; Seattle
Listed on 2026-09-24
-
Software Development
AI Engineer (Applied/Software), DevOps, Software Engineer, Software Architect
Houston;
New York;
San Francisco;
Seattle
Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work.
Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.
Nscale is looking for a
Staff AI Engineer (Specialized)to set technical direction for the inference and reinforcement learning systems at the core of our AI services platform — and for the APIs through which other engineers consume them.
This role owns the architecture of how models are served on Nscale’s GPU cloud, how RL and post-training workloads run on it, and how both are exposed to customers and internal teams as clean, reliable, high-performance interfaces. You’ll work across 2–4 teams spanning serving, post-training, and platform, defining how these systems are built and establishing the standards that create engineering leverage across the organization.
As a Staff engineer, you are the technical authority for this area of the AI stack. Your decisions determine the latency, throughput, and cost profile of every token Nscale serves, and the correctness and efficiency of every RL run on our platform. You resolve ambiguous architectural questions where the answer space is genuinely open — disaggregated versus co-located serving, on-policy versus off-policy RL infrastructure, where the API boundary should sit — and your solutions become the standards others build on.
ResponsibilitiesInference
- Set technical direction for Nscale’s inference serving architecture: request routing, scheduling, continuous batching, KV cache management, prefix caching, and speculative decoding
- Drive the strategy for model efficiency in production — quantization (FP8, INT8/4), sparsity, pruning, distillation, and MoE serving — and the trade-offs between cost, latency, throughput, and model quality
- Lead the resolution of systemic performance and reliability challenges across the serving stack, from kernel-level bottlenecks to fleet-level capacity and multi-tenant isolation
- Own the architecture of Nscale’s RL and post-training systems: RLHF, DPO/GRPO-style methods, reward modelling, and agentic RL with tool calling, off-policy training, and decoupled sampling and policy updates
- Define how inference and training share infrastructure in RL loops — rollout generation, sample buffering, weight synchronization, and the serving engine’s role inside the training system
- Establish standards for fine-tuning services (LoRA, QLoRA, adapters, full fine-tuning) and the data curation and processing workflows that feed them
- Set the design standards for Nscale’s developer-facing APIs, SDKs, and tooling — OpenAI-compatible and native interfaces, OpenAPI 3.x specifications, versioning, and rate limiting — so that inference and RL capabilities are consumable by engineers who never see the underlying systems
- Create reusable frameworks and tooling that multiply the effectiveness of other AI engineers across Nscale
Leadership & direction
- Coach and grow more junior engineers across teams; raise technical capability and engineering quality broadly
- Collaborate with research, product, and infrastructure leadership to align inference and RL platform strategy with customer demand and business direction
- Evaluate emerging serving engines, RL frameworks, and accelerator…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).