More jobs:
AI Systems Performance Specialist
Job in
Durham, Durham County, North Carolina, 27701, USA
Listed on 2026-08-22
Listing for:
Bright Vision Technologies
Full Time
position Listed on 2026-08-22
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer
Job Description & How to Apply Below
AI Systems Performance Specialist
Bright Vision Technologies is seeking a highly experienced AI Systems Performance Specialist with 10+ years of experience in AI infrastructure, machine learning systems, High-Performance Computing (HPC), and performance engineering. The ideal candidate will optimize AI training and inference workloads for maximum performance, scalability, reliability, and cost efficiency. This role requires deep expertise in GPU optimization, distributed training, Large Language Model (LLM) inference, Python, C++, CUDA, and production AI systems, along with the ability to lead performance optimization initiatives across enterprise-scale AI platforms.
Key Responsibilities- Optimize AI training and inference pipelines for maximum throughput, low latency, scalability, and infrastructure efficiency.
- Analyze and improve GPU utilization, memory management, kernel execution, and multi-GPU performance across production AI workloads.
- Design and implement optimization techniques including quantization, pruning, mixed precision, batching, caching, speculative decoding, and model parallelism.
- Profile AI applications using industry-standard performance analysis tools and identify bottlenecks across compute, memory, networking, and storage.
- Optimize distributed training and inference using NCCL, Deep Speed, PyTorch Distributed, Ray, MPI, or similar distributed computing frameworks.
- Collaborate with AI researchers, ML engineers, platform engineers, and infrastructure teams to improve model performance and production reliability.
- Build automated benchmarking frameworks, performance dashboards, monitoring solutions, and regression testing pipelines.
- Evaluate emerging AI hardware, GPU architectures, inference frameworks, and optimization technologies to improve enterprise AI capabilities.
- Drive AI infrastructure cost optimization through efficient resource utilization, cloud optimization, and Fin Ops best practices.
- Mentor engineering teams and provide technical leadership on AI systems architecture, GPU optimization, and performance engineering.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related technical discipline.
- 10+ years of professional experience in performance engineering, AI infrastructure, machine learning systems, High-Performance Computing (HPC), or distributed computing.
- Expert-level programming skills in Python and C++.
- Extensive experience optimizing GPU-accelerated AI workloads using CUDA, distributed training frameworks, and modern deep learning libraries.
- Strong knowledge of Large Language Models (LLMs), deep learning frameworks, model serving, and production AI inference.
- Hands-on experience with profiling tools such as NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, Tensor Board, or similar performance analysis tools.
- Experience deploying and optimizing AI workloads on AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Strong understanding of distributed systems, networking, storage optimization, and AI infrastructure architecture.
- Excellent analytical, troubleshooting, communication, and technical leadership skills.
- Experience optimizing production-scale LLM inference and serving large foundation models.
- Hands-on experience with vLLM, TensorRT-LLM, Deep Speed, Triton Inference Server, CUTLASS, Faster Transformer, or similar AI optimization frameworks.
- Knowledge of model compression, KV cache optimization, speculative decoding, and advanced inference optimization techniques.
- Experience implementing Fin Ops strategies for AI infrastructure cost optimization and resource management.
- Contributions to AI systems research, open-source AI infrastructure projects, patents, or technical publications.
- Familiarity with emerging AI accelerator technologies, including AMD ROCm, Intel oneAPI, or custom AI hardware.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×