×
Register Here to Apply for Jobs or Post Jobs. X

Cloud & Customer Solutions Engineer - DC GPU

Job in Bellevue, King County, Washington, 98009, USA
Listing for: Advanced Micro Devices
Full Time position
Listed on 2026-08-30
Job specializations:
  • Software Development
    AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 150000 - 190000 USD Yearly USD 150000.00 190000.00 YEAR
Job Description & How to Apply Below

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whether you’redesigning next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger— technology that moves the world forward.

Join us and, together, we’ll advance your career.

THE TEAM:

AMD's Data Center GPU organization is transforming the industry with our AI based Graphic Processors. Our primary objective is to design exceptional products that drive the evolution of computing experiences, serving as the cornerstone for enterprise Data Centers, (AI) Artificial Intelligence, HPC and Embedded systems. If this resonates with you, come and joining our Data Center GPU organization where we are building amazing AI powered products with amazing people.

THE ROLE:

As a Cloud and Customer Solutions Engineer on AMD's Applied AI team, you will embed directly with AMD's most strategic AI customers — frontier labs, Neo Cloud providers, CSPs, and AI-native companies — to take AMD Instinct GPU clusters from delivery to sustained production excellence. You own the customer outcome end-to-end: cluster bring-up and certification, workload deployment and performance, production incident response, and the transfer of operational capability that moves customers toward autonomous operation of their AMD fleets.

To be direct about what this role is: despite the "Solutions" title, this is not a pre-sales or demo role. You will write production code, operate live clusters, carry accountability for customer production outcomes, and be the engineer in the room when things break t you learn in the field, you convert into durable improvements — to ROCm, to the open-source serving ecosystem, and to the reference architectures every subsequent deployment inherits.

THE PERSON:

You are a strong production engineer who is energized rather than drained by ambiguity, customer pressure, and environments you do not control. You can debug a distributed training hang at 2am, explain the root cause to a customer VP at 9am, and land the fix upstream by the end of the week. You measure success by customer production outcomes, not code merged or tickets closed.

When something is broken on a cluster you touch, it is your problem until it is fixed or explicitly handed off.

KEY RESPONSIBILITIES:

  • Own customer deployments end-to-end: cluster bring-up and burn-in, production readiness certification, workload onboarding, performance validation, and sustained production operation on AMD Instinct GPU fleets
  • Deploy and tune large-scale training and inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) against customer‑specific workloads and SLOs across cloud, Neo Cloud, and bare‑metal environments
  • Lead root‑cause analysis and resolution of production incidents on customer clusters, including Sev‑1 response, and drive fixes to permanent closure
  • Deploy agentic AI solutions into customer environments in partnership with Agentic Data Engineers, and own their production behavior within the engagement
  • Build the observability, benchmarking, and validation tooling needed to certify clusters as production‑ready and keep them there
  • Transfer operational capability to customer teams — documentation, runbooks, and hands‑on enablement — moving customers up the operator‑autonomy ladder from assisted operation to independent production ownership
  • Contribute field learnings to the Applied AI team's skills library and engagement memory databases, so deployment knowledge compounds across the practice
  • Convert field findings into upstream contributions — ROCm issues and patches, serving‑framework improvements, reference‑architecture updates — and provide structured field signal to AMD product, software, and silicon teams

PREFERRED EXPERIENCE:

  • 5+ years of production software or infrastructure engineering, including significant time operating or deploying systems in environments you did not build (level flexible for…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary