×
Register Here to Apply for Jobs or Post Jobs. X

Technical Program Manager - Cluster Orchestration & Applied Training

Job in Egg Harbor Township, Atlantic County, New Jersey, 08234, USA
Listing for: Coreweave
Apprenticeship/Internship position
Listed on 2026-07-21
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, AI Business & Operations
Salary/Wage Range or Industry Benchmark: 237000 - 261000 USD Yearly USD 237000.00 261000.00 YEAR
Job Description & How to Apply Below
Position: Staff Technical Program Manager - Cluster Orchestration & Applied Training

What You’ll Do

Core Weave is seeking a Staff Technical Program Manager to lead complex, cross‑functional programs across Cluster Orchestration and Applied Training within our AI/ML Platform Services organization.

Cluster Orchestration is the platform layer that makes sure large AI workloads are scheduled, launched, and managed reliably across Core Weave’s clusters.
Applied Training is the layer on top of that infrastructure that helps researchers and customers use it for pre‑training, fine‑tuning, reinforcement learning, evaluations, and sandboxed experimentation.

About the role

In this role, you will partner with engineering, product, infrastructure, and research‑adjacent teams to improve both how workloads run on the cluster and how users interact with the training platform built on top of it. That includes driving programs across orchestration systems such as Slurm‑on‑Kubernetes (SUNK), Kueue, and workflow integrations, while also helping scale the environments, tooling, and operational mechanisms that make training and evaluation workflows easier to use.

This is a highly cross‑functional role for a TPM who combines strong technical depth, excellent execution instincts, and the ability to bring structure and clarity to fast‑moving infrastructure and AI platform initiatives.

The base salary range for this role is $237,000 to $261,000. The starting salary will be determined based on job‑related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

Responsibilities
  • Drive end‑to‑end program execution for cluster orchestration initiatives spanning workload scheduling, self‑service provisioning, upgrade and migration flows, and platform integrations.
  • Lead cross‑functional programs that improve how AI training, evaluation, RL, and mixed workloads run across Core Weave clusters.
  • Partner with engineering and product leaders to define roadmap priorities and deliver measurable improvements in utilization, reliability, scalability, observability, and user experience.
  • Drive delivery for applied training initiatives across pre‑training, fine‑tuning, reinforcement learning, sandbox environments, and evaluation systems.
  • Coordinate dependencies across platform engineering, infrastructure, product, customer‑facing teams, and ecosystem partners to ensure successful launches and clear operational ownership.
  • Build program mechanisms for release readiness, rollout planning, risk management, stakeholder communication, and post‑launch review.
  • Establish success metrics, dashboards, and operating cadences to improve cluster efficiency, workload startup performance, time‑to‑research, and adoption of new platform capabilities.
  • Create clarity across ambiguous technical programs by aligning stakeholders, surfacing tradeoffs early, and driving decisions to resolution.
Who You Are
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 8+ years of technical program management experience in cloud infrastructure, distributed systems, or AI/ML platforms.
  • Experience leading large‑scale cross‑functional programs involving scheduling systems, cluster infrastructure, or ML platform capabilities.
  • Strong technical fluency in Kubernetes, Slurm or comparable schedulers, distributed systems, and AI training workflows.
  • Demonstrated ability to define program metrics and deliver measurable outcomes in performance, reliability, scale, or operational maturity.
  • Excellent communication skills, with experience influencing engineering, product, and executive stakeholders.
Preferred
  • Experience with orchestration and scheduling technologies such as Kubernetes, Slurm, Kueue, Ray, or similar systems.
  • Familiarity with modern AI training and evaluation workflows, including pre‑training, supervised fine‑tuning, reinforcement learning, and experiment or sandbox environments.
  • Understanding of GPU infrastructure, cluster capacity planning, multi‑tenant execution,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary