×
Register Here to Apply for Jobs or Post Jobs. X

Senior GPU Software Engineer

Remote / Online - Candidates ideally in
Hillsboro, Washington County, Oregon, 97104, USA
Listing for: Intel
Remote/Work from Home position
Listed on 2026-08-30
Job specializations:
  • Software Development
    Software Engineer, AI Engineer (Applied/Software), C++ Developer, Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 195000 - 276000 USD Yearly USD 195000.00 276000.00 YEAR
Job Description & How to Apply Below
Position: Senior GPU Performance Software Engineer

The Software and AI (SAI) organization is seeking a highly skilled software engineer to contribute to the development and low-level optimization of

Job Details

Job Description:

About

The Role

The Software and AI (SAI) organization is seeking a highly skilled software engineer to contribute to the development and low-level optimization of oneDNN
, a complex, cross-platform, open-source performance library that serves as the foundation for deep learning applications ().

Please Note

This is a low-level software engineering and hardware-acceleration role. It does not involve building, training, or tuning machine learning models. Instead, you will focus on developing highly optimized math primitives, parallel algorithms, and GPU kernels that power industry-leading AI frameworks (such as OpenVINO, Tensor Flow, PyTorch, and ONNX Runtime) on Intel hardware.

Key Responsibilities Kernel Development and Architecture
  • Develop high-performance GEMM, convolution, and attention kernels for AI workloads
  • Design scalable JIT and codegen infrastructure for GPU kernel generation
Low-Level Optimization
  • Implement fusion and memory-traffic optimizations to maximize hardware utilization
  • Optimize mixed-precision and quantized execution paths (e.g., BF16, FP16, INT8, FP8, FP4, etc.)
Performance Modeling and Profiling
  • Build analytical and empirical performance models for kernel dispatch and tuning
  • Profile and eliminate performance bottlenecks across oneDNN GPU primitives and runtime paths
Hardware and Software Co-Design
  • Co-design GPU primitives and kernel architectures for next-generation Intel GPUs
  • Partner with hardware and compiler teams to shape future accelerator capabilities and software stacks
Infrastructure and Validation
  • Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads
Why Join UsMassive Scale
  • Work on a global, high-impact open-source library that scales AI performance across millions of devices worldwide
Cutting-Edge Hardware
  • Get early access to and influence the software stack for Intel's roadmap of next-generation discrete GPUs
Expert Collaboration
  • Work alongside industry-leading experts in GPU compilers, hardware architecture, and performance libraries
Total Rewards
  • Enjoy a competitive package including stock programs, quarterly bonuses, robust healthcare, and highly flexible hybrid/remote working options
What We're Looking For

To be successful in this role, you should demonstrate the following professional traits:

  • A strong ownership mindset — you take initiative on complex, ambiguous technical problems and drive them to resolution
  • A collaborative approach — you work effectively across hardware, compiler, and framework teams to align on shared technical goals
  • A performance-driven curiosity — you are motivated by squeezing every cycle out of hardware and continuously seek deeper understanding of low-level systems
Qualifications

Minimum Qualifications
  • Education:

    BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a highly technical related field
  • Core Language: 5+ years of professional software development experience with expert-level modern C++
  • Performance Optimizations: 2+ years of hands-on experience in programming and kernel optimization on GPUs (via SYCL/DPC++, OpenCL, CUDA, or HIP), or at least 5+ years of similar low-level performance optimization experience on CPUs
  • Hardware Architecture:
    Strong foundations in computer architecture, cache hierarchies, memory subsystems, and parallel programming paradigms (e.g., multi-threading, SIMD/vectorization)
Preferred Qualifications
  • Math Libraries:
    Experience developing high-performance math libraries (e.g., GEMM, convolution, reduction, or FFT kernels)
  • Low-Level Tuning:
    Hands-on experience with GPU assembly-level tuning or compiler optimization
  • Parallel APIs:
    Familiarity with parallel programming APIs such as OpenMP or oneTBB
  • AI Workload Context:
    Basic understanding of deep learning primitives (e.g., forward/backward passes) to understand how library code is utilized by upstream frameworks
Job Type

Experienced Hire

Shift

Shift 1 (United States of America)

Primary Location

US, Oregon, Hillsboro

Additional Locations

US, California,…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary