Senior Software Engineer - AI Inference
Listed on 2026-09-03
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Python, Software Engineer
Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.
Requirements: Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA. A degree in Computer Science or Computer Engineering is required along with deep knowledge of GPU architecture and profiling tools.
Key
Skills:
CUDA, Python, C++, Rust, LLM Inference, VLM Inference, GPU Architecture, Nsight Systems, Nsight Compute, PyTorch Profiler, TensorRT-LLM, vLLM, Triton, CUTLASS, Distributed Systems, Quantization
Benefits: Equity, Generous Benefits Package
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).