×
Register Here to Apply for Jobs or Post Jobs. X

AI Performance Engineer – GPU & ROCm

Job in Toronto, Ontario, C6A, Canada
Listing for: Sapphire Stream Technology
Full Time position
Listed on 2026-08-03
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 90000 - 140000 CAD Yearly CAD 90000.00 140000.00 YEAR
Job Description & How to Apply Below

We are looking for a GPU AI Performance Engineer to optimize and deploy AI and machine-learning workloads using ROCm.

You will improve the speed, memory efficiency, and scalability of AI training and inference workloads. You will also work with GPU software, AI frameworks, profiling tools, and cloud or containerized environments.

Key Responsibilities
  • Improve training speed, inference latency, memory usage, and multi-GPU scalability.
  • Work with large language models, vision models, multimodal AI, and generative AI.
  • Develop and troubleshoot applications using ROCm and HIP.
  • Work with PyTorch, Tensor Flow, ONNX Runtime, vLLM, SGLang, or MIGraph

    X.
  • Profile GPU applications and identify performance bottlenecks.
  • Apply mixed precision, quantization, kernel tuning, operator fusion, and memory optimization.
  • Support AI deployment on Linux, cloud, edge, and containerized platforms.
  • Develop performance benchmarks and automated validation tests.
  • Collaborate with hardware, compiler, runtime, framework, and infrastructure teams.
Required Qualifications
  • Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Strong programming skills in C++ and Python.
  • Experience with Linux and GPU computing.
  • Hands‑on experience with ROCm, HIP, CUDA, or similar GPU technologies.
  • Understanding of GPU architecture, parallel programming, and memory management.
  • Experience with PyTorch, Tensor Flow, ONNX Runtime, or similar AI frameworks.
  • Experience profiling and optimizing GPU applications.
  • Knowledge of AI model training, inference, and performance optimization.
Preferred Qualifications
  • Experience optimizing large language models or generative AI applications.
  • Experience with vLLM, SGLang, or distributed inference.
  • Familiarity with LLVM, MLIR, or compiler technologies.
  • Experience migrating workloads between CUDA and HIP.
  • Knowledge of Kubernetes, containers, and distributed AI systems.
  • Contributions to ROCm or other open-source GPU projects.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary