CUDA Developer
Listed on 2026-08-15
-
Software Development
Software Engineer, AI Engineer (Applied/Software), Computer Software / Middleware
CUDA Developer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job TitleCUDA Developer
Location100% Remote (U.S.)
Position TypeFull-time, Direct W2
Salary Range$85,000–$110,000 Annually
Experience Required6+ years
SponsorshipU.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job SummaryWe are seeking a CUDA Developer with deep expertise in CUDA programming, GPU architecture, and high-performance computing to design and optimize compute-intensive workloads on modern accelerator hardware. This role focuses on extracting maximum performance from GPU platforms for AI training, inference, scientific computing, and high-throughput data processing workloads. The ideal candidate combines low-level systems mastery with strong software engineering practices, and has a track record of delivering measurable performance improvements on production GPU systems.
In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field.
- Six or more years of experience in GPU programming and performance engineering.
- Deep expertise in CUDA C/C++ and GPU programming models.
- Strong understanding of modern GPU architectures, memory hierarchies, and execution models.
- Hands-on experience profiling and optimizing GPU workloads in production.
- Familiarity with NCCL, MPI, and high-performance interconnect technologies.
- Experience integrating custom kernels into ML frameworks.
- Strong C++ skills and familiarity with modern systems programming practices.
- Solid grounding in linear algebra and numerical methods.
- Strong communication and collaboration skills with research and engineering teams.
- Experience with Triton, CUTLASS, or other GPU kernel authoring frameworks.
- Familiarity with TensorRT, Faster Transformer, or vLLM internals.
- Exposure to compiler infrastructure such as LLVM or MLIR.
- Open-source contributions to GPU or ML performance libraries.
- Experience with large-scale distributed training infrastructure.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).