More jobs:
GPUaaS Kubernetes Platform Engineer
Job in
Irving, Dallas County, Texas, 75084, USA
Listed on 2026-07-13
Listing for:
Veriipro
Full Time
position Listed on 2026-07-13
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure, Systems Engineer
Job Description & How to Apply Below
We are seeking a skilled GPUaaS Kubernetes Platform Engineer to design, operate, and support scalable GPU-enabled cloud infrastructure. The role focuses on managing Kubernetes and Open Shift platforms, enabling AI/ML workloads, optimizing GPU resource utilization, and ensuring the reliability and performance of high-performance computing environments.
Roles and Responsibilities- Operate and maintain Kubernetes and Open Shift GPU platforms supporting AI, ML, and high-performance computing workloads.
- Configure, deploy, and manage GPU-enabled infrastructure, including accelerator resources, GPU scheduling, and workload optimization.
- Enable and support AI/ML workloads through scalable GPU-as-a-Service (GPUaaS) platforms.
- Implement and maintain CI/CD pipelines for GPU-based applications and infrastructure deployments.
- Manage platform scalability, availability, and performance to ensure reliable GPU service delivery.
- Perform Kubernetes/Open Shift administration, including cluster configuration, upgrades, monitoring, and troubleshooting.
- Develop automation for GPU platform operations, provisioning, and lifecycle management.
- Monitor GPU utilization, capacity planning, and resource optimization across environments.
- Create and maintain operational dashboards, scaling reports, and platform performance metrics.
- Develop and maintain technical documentation, runbooks, troubleshooting guides, and operational procedures.
- Collaborate with AI/ML teams, infrastructure teams, and Dev Ops engineers to deliver enterprise-grade GPU services.
- Strong experience with Kubernetes administration and container orchestration.
- Hands-on experience managing Open Shift environments.
- Experience operating GPU-enabled Kubernetes platforms and accelerator workloads.
- Knowledge of GPU infrastructure concepts, scheduling, and resource management.
- Experience with CI/CD tools and Dev Ops practices.
- Strong understanding of Linux systems, networking, storage, and cloud infrastructure.
- Experience with monitoring, troubleshooting, and performance optimization of distributed platforms.
- Ability to create operational documentation, runbooks, and support procedures.
- Experience with AI/ML infrastructure platforms and MLOps environments.
- Knowledge of GPU technologies such as NVIDIA GPU architectures, CUDA, and GPU operators.
- Experience with infrastructure automation tools such as Terraform, Ansible, or similar technologies.
- Familiarity with cloud platforms and hybrid infrastructure environments.
- Experience supporting large-scale production Kubernetes platforms.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×