×
Register Here to Apply for Jobs or Post Jobs. X

Technical Program Manager III, GPU Infrastructure Reliability, Cloud

Job in Sunnyvale, Santa Clara County, California, 94087, USA
Listing for: Google
Full Time position
Listed on 2026-06-21
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, IT Project Manager, Systems Engineer
Salary/Wage Range or Industry Benchmark: 200000 - 250000 USD Yearly USD 200000.00 250000.00 YEAR
Job Description & How to Apply Below
Position: Technical Program Manager III, GPU Infrastructure Reliability, Google Cloud

Benefits

In accordance with Washington state law, we are highlighting our comprehensive benefits package, which is available to all eligible US based employees. Benefits for this role include:

  • Health, dental, vision, life, disability insurance
  • Retirement Benefits: 401(k) with company match
  • Paid Time Off: 20 days of vacation per year, accruing at a rate of 6.15 hours per pay period for the first five years of employment
  • Sick Time: 40 hours/year (increased to 69 hours/year for Seattle) including 5 discretionary sick days per instance
  • Maternity Leave (Short-Term Disability + Baby Bonding): 28-30 weeks
  • Baby Bonding Leave: 18 weeks
  • Holidays: 13 paid days per year

Note:

By applying to this position you will have an opportunity to share your preferred working location from the following:
Sunnyvale, CA, USA;
Kirkland, WA, USA

.

Minimum Qualifications
  • Bachelor's degree in a technical field, or equivalent practical experience.
  • 5 years of experience in program management.
  • Experience with infrastructure reliability.
  • Experience with GPUs or GPU Systems.
Preferred Qualifications
  • 5 years of experience managing cross-functional or cross-team projects.
  • 5 years of experience in technical program management, with a focus on software engineering and ML infrastructure projects.
  • Knowledge of software development, distributed systems, and ML infrastructure or GPU systems.
  • Ability to think critically and solve problems.
  • Excellent project management skills, and experience with project planning, execution, and risk management.
  • Excellent communication and collaboration skills, with the ability to build relationships and influence across all levels of the organization.
About the Job

A problem isn’t truly solved until it’s solved for all. That’s why Googlers build products that help create opportunities for everyone, whether down the street or across the globe. As a Technical Program Manager at Google, you’ll use your technical expertise to lead complex, multi-disciplinary projects from start to finish. You’ll work with stakeholders to plan requirements, identify risks, manage project schedules, and communicate clearly with cross‑functional partners across the company.

You’ll be equally comfortable explaining your team’s analyses and recommendations to executives as you are discussing the technical tradeoffs in product development with engineers.

To empower AI innovation by accelerating the delivery, cloud‑based accelerator (GPU) NPI’s built into large‑scale supercomputer clusters, including next‑gen cross‑functional development, customer and vendor partnerships, and ML workload monitoring and diagnostic tooling.

As a GPU Technical Program Manager for Google Cloud’s AI and Computing Infrastructure team, you will be at the forefront of AI innovation, leading the end‑to‑end development and delivery of next‑generation Cloud GPU products from initial concept to full‑scale production. You will take charge of software qualification and release strategies for AI hyper‑compute clusters, collaborating deeply with engineering, product, and capacity planning teams to align customer and business priorities.

Beyond managing critical escalations and mitigating risks, this is a unique opportunity to shape cross‑functional initiatives alongside Application Centric Infrastructure (ACI) leadership and Technical Program Managers (TPMs) across the broader organization to streamline customer onboarding and scaled support for our largest, most complex Cloud ML solutions.

The ML, Systems, & Cloud AI (MSCA) organization at Google designs, implements, and manages the hardware, software, machine learning, and systems infrastructure for all Google services (Search, You Tube, etc.) and Google Cloud. Our end users are Googlers, Cloud customers and the billions of people who use Google services around the world.

We prioritize security, efficiency, and reliability across everything we do— from developing our latest TPUs to running a global network, while driving towards shaping the future of hyperscale computing. Our global impact spans software and hardware, including Google Cloud’s Vertex AI, the leading AI platform for bringing Gemini models to enterprise customers.

Indivi…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary