×
Register Here to Apply for Jobs or Post Jobs. X

Devops Platform Engineer

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: Advanced Micro Devices
Full Time position
Listed on 2026-09-10
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 190000 USD Yearly USD 150000.00 190000.00 YEAR
Job Description & How to Apply Below

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whether you’redesigning next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger— technology that moves the world forward.

Join us and, together, we’ll advance your career.

THE ROLE:

We are seeking a hands-on Platform Engineer to build and operate the infrastructure foundations that enable engineering teams to deliver reliable, scalable, and secure products. This role will shape the internal developer platform across Kubernetes, observability, AI-enabled workloads, CI/CD, and infrastructure as code.

You will work closely with software, Dev Ops, security, and AI/ML teams to turn platform capabilities into a simple and dependable developer experience. The ideal candidate is comfortable moving between architecture decisions, production operations, automation, and practical enablement of partner teams.

KEY RESPONSIBILITIES:
  • Design, build, and operate scalable Kubernetes-based platform capabilities for development, test, and production workloads.
  • Create reusable infrastructure-as-code modules, patterns, and automated workflows using tools such as Terraform, Ansible, Helm, Kustomize, or equivalent technologies.
  • Develop and evolve platform observability: metrics, logs, traces, dashboards, alerting, service-level objectives, and operational runbooks.
  • Improve platform reliability, capacity management, security posture, and cost efficiency through automation and data-driven operational practices.
  • Build the infrastructure required to deploy, operate, observe, and govern AI/ML workloads, including model-serving, GPU-enabled compute, data-access, and workload-isolation patterns where applicable.
  • Enable self-service for product and engineering teams through well-designed platform APIs, templates, documentation, golden paths, and CI/CD integrations.
  • Partner with application, security, infrastructure, and AI teams to define platform standards and remove delivery bottlenecks.
  • Investigate complex production issues, lead root-cause analysis, and implement durable preventative improvements.
  • Contribute to technical roadmaps, architecture reviews, and engineering standards for the platform.
PREFERRED EXPERIENCE:
  • Experience designing internal developer platforms or large-scale shared infrastructure.
  • Experience with Prometheus, Python, Bash, Grafana, Open Telemetry, Loki, Tempo, Elastic, Datadog, Splunk, or comparable observability systems.
  • Experience with IBM LSF and Netapp Storage
  • Experience with Git Ops operating models.
  • Experience supporting AI/ML infrastructure, model serving, GPU scheduling, Kubernetes operators, or AI workload orchestration.
  • Knowledge of platform security practices, including identity and access management, secrets management, policy enforcement, image security, and supply-chain security.
  • Experience with service mesh, API gateways, ingress, networking, or multi-cluster Kubernetes architectures.
  • Demonstrated ownership of production systems and a bias toward automation, operational excellence, and continuous improvement.
  • Semiconductor experience very helpful, EDA experience a plus.
Required Qualifications
  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
  • 6+ years of experience in software engineering, Dev Ops, SRE, cloud infrastructure, or platform engineering.
  • Strong hands-on experience operating and automating Kubernetes environments.
  • Experience with container technologies and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary