GPU Systems Engineer
Plymouth, Hennepin County, Minnesota, USA
Listed on 2026-10-04
-
Software Development
AI Engineer (Applied/Software)
-Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: GPU Systems Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience
Required:
6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EADHolders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
We are seeking a GPU Systems Engineer with deepexpertisein CUDA programming, GPU architecture, and high-performance computing to design andoptimizecompute-intensive workloads on modern accelerator hardware. This role focuses on extracting maximum performance from GPU platforms for AI training, inference, scientific computing, and high-throughput data processing workloads. The ideal candidate combines low-level systems mastery with strong software engineering practices, andhasa track recordof delivering measurable performance improvements on production GPU systems.
In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, andwill be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, anda track recordof shipping meaningful work that holds up well in production.
- Design and implement high-performance CUDA kernels for compute-intensive workloads across AI and HPC use cases.
- Profile and optimize GPU code using tools such as Nsight Systems, Nsight Compute, and CUDA profilers.
- Tune memory access patterns, occupancy, register usage, and shared memoryutilizationfor peak performance.
- Develop highly optimized libraries for linear algebra, attention, and other ML primitives.
- Optimize multi-GPU and multi-node training using NCCL, RDMA, and high-performance networking.
- Implement custom operators and fused kernels inPyTorch, JAX, or Triton.
- Collaborate with ML engineers toidentifyperformance bottlenecks in training and inference pipelines.
- Develop benchmarks and regression tests to safeguard performance over time.
- Evaluate new GPU architectures and feature sets, andadvise on adoption strategy.
- Contribute to compiler-level optimizations for tensor programs where appropriate, working at the boundary between ML frameworks and underlying acceleratorcodegento unlock performance not reachable through framework-level tuning alone.
- Optimize memory hierarchy usage across HBM, L2, shared memory, and registers.
- Implement mixed-precision and quantized compute paths that maximize accelerator throughput while preserving numerical fidelity within bounds acceptable for the target workloads.
- Document performance characteristics, design decisions, and tuning playbooks for internal teams.
- Stay current with GPU architecture, CUDA evolution, and emerging accelerator technologies.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, ora related field.
- Six or more years of experience in GPU programming and performance engineering.
- Deepexpertisein CUDA C/C++ and GPU programming models.
- Strong understanding of modern GPU architectures, memory hierarchies, and execution models.
- Hands-on experience profiling andoptimizingGPU workloads in production.
- Familiarity with NCCL, MPI, and high-performance interconnect technologies.
- Experience integrating…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).