×
Register Here to Apply for Jobs or Post Jobs. X

AI Product Engineer Houston San Francisco; Seattle

Job in New York, New York County, New York, 10261, USA
Listing for: Nscale
Full Time position
Listed on 2026-09-24
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), DevOps, Software Engineer, Software Architect
Salary/Wage Range or Industry Benchmark: 220000 - 293333 USD Yearly USD 220000.00 293333.00 YEAR
Job Description & How to Apply Below

Houston;
New York;
San Francisco;
Seattle

About Nscale

Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work.

Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.

About the Role

Nscale is looking for a
Staff AI Engineer (Specialized)to set technical direction for the inference and reinforcement learning systems at the core of our AI services platform — and for the APIs through which other engineers consume them.

This role owns the architecture of how models are served on Nscale’s GPU cloud, how RL and post-training workloads run on it, and how both are exposed to customers and internal teams as clean, reliable, high-performance interfaces. You’ll work across 2–4 teams spanning serving, post-training, and platform, defining how these systems are built and establishing the standards that create engineering leverage across the organization.

As a Staff engineer, you are the technical authority for this area of the AI stack. Your decisions determine the latency, throughput, and cost profile of every token Nscale serves, and the correctness and efficiency of every RL run on our platform. You resolve ambiguous architectural questions where the answer space is genuinely open — disaggregated versus co-located serving, on-policy versus off-policy RL infrastructure, where the API boundary should sit — and your solutions become the standards others build on.

Responsibilities

Inference

  • Set technical direction for Nscale’s inference serving architecture: request routing, scheduling, continuous batching, KV cache management, prefix caching, and speculative decoding
  • Drive the strategy for model efficiency in production — quantization (FP8, INT8/4), sparsity, pruning, distillation, and MoE serving — and the trade-offs between cost, latency, throughput, and model quality
  • Lead the resolution of systemic performance and reliability challenges across the serving stack, from kernel-level bottlenecks to fleet-level capacity and multi-tenant isolation
  • Own the architecture of Nscale’s RL and post-training systems: RLHF, DPO/GRPO-style methods, reward modelling, and agentic RL with tool calling, off-policy training, and decoupled sampling and policy updates
  • Define how inference and training share infrastructure in RL loops — rollout generation, sample buffering, weight synchronization, and the serving engine’s role inside the training system
  • Establish standards for fine-tuning services (LoRA, QLoRA, adapters, full fine-tuning) and the data curation and processing workflows that feed them
  • Set the design standards for Nscale’s developer-facing APIs, SDKs, and tooling — OpenAI-compatible and native interfaces, OpenAPI 3.x specifications, versioning, and rate limiting — so that inference and RL capabilities are consumable by engineers who never see the underlying systems
  • Create reusable frameworks and tooling that multiply the effectiveness of other AI engineers across Nscale

Leadership & direction

  • Coach and grow more junior engineers across teams; raise technical capability and engineering quality broadly
  • Collaborate with research, product, and infrastructure leadership to align inference and RL platform strategy with customer demand and business direction
  • Evaluate emerging serving engines, RL frameworks, and accelerator…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary