×
Register Here to Apply for Jobs or Post Jobs. X

Remote | AWS Trainium Kernel Engineer; NKI

Remote / Online - Candidates ideally in
New York, New York County, New York, 10261, USA
Listing for: 24-Mag Llc
Part Time, Remote/Work from Home position
Listed on 2026-09-10
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 60 - 80 USD Hourly USD 60.00 80.00 HOUR
Job Description & How to Apply Below
Position: Remote | AWS Trainium Kernel Engineer (NKI) — $60–$80/hour
Location: New York

We are sharing a specialised part-time consulting opportunity for experienced kernel engineers with hands‑on expertise in the Neuron Kernel Interface (NKI), AWS Trainium/Inferentia2 hardware, low‑level performance optimisation, and CUDA‑to‑NKI migration.

This role focuses on evaluating NKI kernel‑development tasks for technical correctness, hardware appropriateness, numerical fidelity, and performance quality. Selected experts will review Trainium‑specific implementations, migration decisions, memory‑management strategies, profiling results, and cross‑platform numerical behaviour while providing clear, rubric‑based technical feedback.

Key Responsibilities NKI Kernel Development Review
  • Evaluate kernels developed using the Neuron Kernel Interface (NKI)
  • Assess whether implementations appropriately target AWS Trainium and Inferentia2 hardware
  • Review low‑level computation patterns for technical correctness
  • Identify inefficient, incorrect, or hardware‑inappropriate implementation choices
  • Apply practical judgement grounded in hands‑on NKI development experience
CUDA‑to‑NKI Migration
  • Review migrations of existing CUDA kernels to NKI
  • Assess whether computational semantics are preserved across platforms
  • Identify translation errors, unsupported assumptions, or inefficient migration strategies
  • Evaluate whether NKI implementations appropriately account for Trainium architecture
  • Distinguish faithful migrations from implementations that merely reproduce surface‑level CUDA structure
Tile‑Based Computation
  • Assess tile decomposition and computation strategies
  • Review partitioning decisions against NKI execution constraints
  • Evaluate whether kernels make effective use of available compute resources
  • Identify inefficient tiling or data‑movement patterns
  • Assess whether implementation choices align with NKI programming requirements
Memory Hierarchy Management
  • Review use of SBUF, PSUM, and HBM
  • Assess data placement and movement across Trainium memory hierarchies
  • Evaluate memory‑bandwidth utilisation and locality
  • Identify unnecessary transfers or memory bottlenecks
  • Review implementation decisions affecting on‑chip memory efficiency
DMA & Data Movement
  • Evaluate DMA orchestration within NKI kernels
  • Review sequencing of computation and data‑transfer operations
  • Identify stalls, inefficient transfer patterns, or synchronisation issues
  • Assess whether data movement appropriately overlaps with computation
  • Evaluate implementation choices affecting pipeline utilisation
Trainium Performance Optimisation
  • Review Trainium‑specific optimisation strategies
  • Assess Neuron Core pipeline utilisation, tensor‑engine throughput, and memory‑bandwidth behaviour
  • Identify performance bottlenecks within kernel implementations
  • Evaluate whether optimisation decisions are supported by profiling evidence
  • Review trade‑offs affecting throughput, latency, and resource utilisation
Numerical Correctness
  • Evaluate numerical consistency between GPU and Trainium implementations
  • Review differences caused by accumulation order, rounding behaviour, and mixed‑precision semantics
  • Assess appropriate tolerances for cross‑platform comparisons
  • Identify numerical discrepancies that indicate implementation defects
  • Distinguish expected hardware‑level variation from substantive correctness problems
Precision & Data Types
  • Review kernels using supported formats such as FP32, BF16, FP8, and INT8
  • Assess precision choices against computational requirements
  • Evaluate mixed‑precision behaviour and numerical stability
  • Identify inappropriate casting or accumulation strategies
  • Review whether performance gains are achieved without compromising required correctness
AWS Neuron Ecosystem
  • Evaluate implementations using the AWS Neuron SDK
  • Review interactions between kernel code, compilation, and Trainium execution
  • Assess compiler‑related behaviours where relevant
  • Apply familiarity with NKI kernel libraries and Neuron tooling
  • Identify implementation issues arising from platform‑specific constraints
Benchmarking & Validation
  • Review benchmark results for Trainium workloads
  • Assess performance comparisons and experimental methodology
  • Evaluate workloads running on Trn1 or Trn2 instances where applicable
  • Determine whether claimed performance…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary