×
Register Here to Apply for Jobs or Post Jobs. X

On-Premise LLM Inference & GPU Systems Engineer

Job in Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listing for: Compunnel, Inc.
Full Time position
Listed on 2026-08-22
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 120000 - 150000 USD Yearly USD 120000.00 150000.00 YEAR
Job Description & How to Apply Below

North Carolina, Charlotte

06/05/2026

Contract

Active

Job Description:

Job Summary

We are seeking an On-Premise LLM Inference & GPU Systems Engineer to build, optimize, and support a large-scale enterprise Generative AI infrastructure environment. This role is focused exclusively on Large Language Model (LLM) inference operations within a private on-premises ecosystem utilizing NVIDIA H200 GPU clusters and Open Shift AI. The ideal candidate will possess deep expertise in GPU runtime optimization, inference serving platforms, Kubernetes-based orchestration, and production-scale deployment of open-source LLMs.

This position will be responsible for maximizing inference performance, operational efficiency, and platform reliability across enterprise AI workloads.

Key Responsibilities
  • Design, deploy, and maintain large-scale on-premises LLM inference infrastructure supporting enterprise Generative AI workloads.
  • Optimize runtime performance of token generation pipelines, including prefill/decode optimization and KV cache management.
  • Deploy, configure, and manage inference serving platforms such as vLLM and TensorRT-LLM.
  • Optimize GPU utilization, throughput, batching strategies, latency, and resource efficiency across production inference environments.
  • Manage workload scheduling and orchestration using Kubernetes-based GPU orchestration platforms and RunAI.
  • Oversee the complete lifecycle of open-source language models, including onboarding, deployment, version management, monitoring, and retirement.
  • Manage and support enterprise Hugging Face model deployment workflows and operational processes.
  • Operate, maintain, and optimize the Open Shift AI ecosystem supporting Generative AI applications and services.
  • Monitor platform performance, identify bottlenecks, and implement optimization strategies to improve inference efficiency and scalability.
  • Collaborate with AI, platform engineering, infrastructure, and operations teams to ensure reliable service delivery.
  • Implement operational best practices related to platform availability, monitoring, security, and governance.
  • Develop automation, deployment processes, and operational procedures to support platform scalability and maintainability.
  • Troubleshoot and resolve infrastructure, inference, performance, and deployment issues across the AI ecosystem.
  • Create and maintain technical documentation, operational runbooks, and platform standards.
Required Qualifications
  • 5+ years of experience as an LLM Systems Engineer, AI Infrastructure Engineer, AI Platform Engineer, or related role.
  • 5+ years of hands‑on experience supporting NVIDIA GPU environments and runtime optimization techniques.
  • Experience optimizing token generation pipelines, including KV cache management and prefill/decode optimization strategies.
  • Strong experience deploying and managing inference frameworks such as vLLM and TensorRT-LLM.
  • 3+ years of experience with Open Shift AI and containerized AI platform operations.
  • 3+ years of experience with GPU orchestration technologies, including RunAI and Kubernetes-based environments.
  • Experience deploying, managing, and supporting open-source Large Language Models in production environments.
  • Proven experience managing the Hugging Face model lifecycle, including onboarding, deployment, version management, and retirement.
  • Strong understanding of AI inference architectures, GPU resource management, workload optimization, and performance tuning.
  • Experience working with containerization technologies, Kubernetes, and cloud-native application platforms.
  • Strong troubleshooting, performance analysis, and problem‑solving skills.
  • Excellent communication and collaboration skills with the ability to work across infrastructure, platform, and AI engineering teams.
Preferred Qualifications
  • Experience supporting enterprise-scale Generative AI platforms and private AI infrastructure environments.
  • Experience optimizing large-scale LLM inference workloads in highly regulated or secure environments.
  • Knowledge of AI observability, monitoring, logging, and performance analytics tools.
  • Experience implementing infrastructure automation and operational tooling for AI platforms.
  • Familiarity with enterprise governance, security, and compliance practices for AI workloads.
  • Experience supporting multi-cluster Kubernetes or Open Shift environments.
  • Knowledge of emerging trends and best practices in LLM inference optimization and AI platform engineering.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary