CUDA/C++ Performance Engineer Differentiable Physics Simulator
Job in
San Jose, Santa Clara County, California, 95199, USA
Listed on 2026-07-02
Listing for:
Honda Research Institute USA
Full Time
position Listed on 2026-07-02
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Software Engineer, Computer Software / Middleware
Job Description & How to Apply Below
CUDA/C++ Performance Engineer for Differentiable Physics Simulator - Honda Research Institute USA
Job Number: P25 T05
Location:
San Jose, CA
Honda Research Institute USA (HRI‑US) seeks a self‑motivated engineer to join our Intelligent Robotics Research division. The role focuses on improving the performance of a CUDA/C++ differentiable physics simulator across the GPU backend, including CUDA kernels, host/device data flow, sparse solver structure, and backward workflows used in optimization. Responsibilities include profiling forward and backward simulation workloads, identifying bottlenecks, improving GPU utilization, and reducing CPU/GPU synchronization and transfer overhead.
Key Responsibilities- Profile bottlenecks across CUDA kernels, memory transfers, synchronization, sparse assembly, solver steps, and differentiable rollout paths.
- Determine the highest‑impact performance lever for each bottleneck, whether kernel tuning, data residency, batching, stream usage, solver changes, or reduction/assembly redesign.
- Improve existing CUDA backend architecture, including host/device data flow, CUDA kernels, sparse assembly, and solver structure.
- Evaluate trade‑offs between targeted optimization, architectural refactoring, and larger rewrites when justified by profiling evidence.
- Measure and validate improvements in speed, GPU utilization, correctness, and numerical reproducibility.
- Deliver results in accordance with project timelines.
- Prepare written and oral technical reports and demonstrations.
- Collaborate with teams of scientists and engineers in Honda’s regional and global R&D offices, communicating profiling results, trade‑offs, and implementation outcomes to audiences with varying CUDA experience.
- Strong C++ and CUDA C++ experience in production or research codebases.
- Proven experience profiling and optimizing CUDA kernels with tools such as NVIDIA Nsight Systems, Nsight Compute, or equivalent GPU profiling workflows.
- Comfortable editing low‑level GPU code involving reductions, atomics, sparse matrices, memory coalescing, launch configuration, and synchronization.
- Experience reducing CPU/GPU transfer overhead using better data residency, batching, pinned memory, async copies, streams, or kernel fusion.
- Familiarity with numerical simulation, optimization, or differentiable physics workflows.
- At least 1 year of hands‑on experience with the qualifications above.
- Experience with contact‑rich physics simulation.
- Experience with differentiable simulation or trajectory optimization.
- Experience optimizing sparse linear solvers on GPU.
- Experience tuning CUDA MPS workloads or multi‑process GPU scheduling.
- 3+ yearsof hands‑on experience with the qualifications above preferred.
8/17/2026
Contract Duration1 year
Position KeywordsCUDA, Differentiable Simulators
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×