×
Register Here to Apply for Jobs or Post Jobs. X

Remote | GPU Kernel Engineer

Remote / Online - Candidates ideally in
New York, New York County, New York, 10261, USA
Listing for: 24-Mag Llc
Full Time, Part Time, Remote/Work from Home position
Listed on 2026-09-10
Job specializations:
  • Software Development
    Software Testing, AI Reliability/ Performance Engineer, Software Engineer, AI QA / Validation Engineer
Salary/Wage Range or Industry Benchmark: 60 - 80 USD Hourly USD 60.00 80.00 HOUR
Job Description & How to Apply Below
Position: Remote | GPU Kernel Engineer — $60–$80/hour
Location: New York

We are sharing a specialised part-time consulting opportunity for experienced GPU and accelerator engineers with hands-on expertise in kernel development, numerical validation, performance optimisation, profiling, and low-level compute frameworks.

This role focuses on evaluating GPU and accelerator kernel-development tasks for correctness, completeness, reproducibility, and performance quality. Selected engineers will review kernel implementations, benchmark methodology, compilation and runtime behaviour, numerical tolerances, and optimisation decisions across multiple accelerator ecosystems.

Key Responsibilities
GPU Kernel Development Review
  • Evaluate GPU and accelerator kernels for technical correctness and completeness
  • Review implementations developed from specifications or reference operators
  • Assess whether kernels correctly implement intended mathematical behaviour
  • Identify implementation defects, unsupported assumptions, or incomplete solutions
  • Apply practical judgement grounded in hands-on kernel engineering experience
CUDA & Accelerator Frameworks
  • Review kernel development across frameworks such as CUDA, Triton, NKI, and Pallas (JAX)
  • Assess framework-specific implementation choices and execution constraints
  • Evaluate kernel translations or migrations between different frameworks
  • Identify incorrect assumptions when moving implementations across accelerator ecosystems
  • Compare alternative kernel implementations for correctness and technical quality
Numerical Correctness
  • Assess kernel outputs against suitable reference implementations
  • Evaluate absolute, relative, and ULP-based tolerances
  • Determine whether numerical differences fall within acceptable limits
  • Review floating-point behaviour and precision-related edge cases
  • Identify discrepancies caused by implementation defects rather than expected numerical variation
Performance Profiling & Benchmarking
  • Evaluate kernel performance using tools such as Nsight, Nsight Compute (ncu), roofline analysis, or framework-native profilers
  • Assess whether benchmark methodology produces fair and meaningful comparisons
  • Review latency, throughput, utilisation, and memory behaviour
  • Identify misleading benchmarking practices or inappropriate baselines
  • Determine whether claimed performance improvements are supported by evidence
Kernel Performance Optimisation
  • Review optimisation strategies for compute and memory efficiency
  • Assess tiling, vectorisation, parallelisation, and workload decomposition
  • Evaluate trade-offs between arithmetic throughput and memory movement
  • Identify bottlenecks affecting kernel performance
  • Assess whether optimisations preserve numerical correctness
Memory Hierarchy Optimisation
  • Review use of registers, shared memory, caches, and accelerator-specific memory resources
  • Evaluate shared-memory tiling, register pressure, bank conflicts, and coalescing patterns
  • Identify inefficient memory-access behaviour
  • Assess data locality and memory-bandwidth utilisation
  • Evaluate whether memory optimisations appropriately match the target hardware
Compilation & Runtime Validation
  • Diagnose common kernel compilation and runtime failures
  • Review issues involving driver incompatibilities, out-of-memory conditions, launch configurations, shape or stride mismatches, and autotuning failures
  • Determine whether failures originate from kernel logic, environment configuration, or runtime assumptions
  • Evaluate proposed debugging approaches and corrective actions
  • Assess whether tasks execute reliably in their intended environment
Kernel Translation & Hardware Migration
  • Review kernels translated or lowered across programming frameworks
  • Evaluate migration between different accelerator targets
  • Assess whether computational semantics and performance assumptions remain valid
  • Identify platform-specific behaviour that requires redesign rather than direct translation
  • Evaluate migration quality across GPU and custom-accelerator environments
Debugging & Operator Fusion
  • Review debugging tasks involving incorrect or unstable kernel implementations
  • Diagnose failures using outputs, profiler data, runtime behaviour, and source code
  • Evaluate operator-fusion strategies where relevant
  • Assess whether fused kernels preserve intended semantics
  • Identify optimisation decisions that introduce correctness or maintainability issues
Compiler & Lowering Concepts
  • Evaluate kernel tasks involving compiler or intermediate-representation concepts where applicable
  • Review transformations between high-level operators and accelerator-level implementations
  • Assess lowering decisions for…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary