×
Register Here to Apply for Jobs or Post Jobs. X

Senior​/ML Compiler Engineer

Job in San Jose, Santa Clara County, California, 95112, USA
Listing for: Step Up Recruiting
Full Time position
Listed on 2026-07-27
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Software Engineer, Computer Software / Middleware, C++ Developer
Job Description & How to Apply Below
Position: Senior / Staff ML Compiler Engineer

Senior / Staff ML Compiler Engineer

Location:

San Jose or Irvine, CA | Full-Time

Our client is a fast-growing fabless semiconductor company building next-generation, energy-efficient domain-specific processors designed for edge AI, wireless communications, advanced radar, computer vision, and autonomous systems.

Founded by industry veterans, our client is developing breakthrough System-on-Chip (SoC) architectures leveraging RISC-V and advanced compute technologies to deliver real-time intelligence at the sensor edge.

This is an opportunity to join a highly innovative engineering team working at the intersection of semiconductor architecture, AI acceleration, and compiler technology.

Position Summary

We are seeking a Senior / Staff ML Compiler Engineer to develop and optimize compiler technologies that unlock the performance of next-generation custom silicon.

This is not a traditional application software engineering role.

The ideal candidate brings deep expertise in compiler backend development, code generation, runtime optimization, hardware/software co-design, and performance tuning for compute-intensive architectures.

This individual will work closely with architecture, silicon, systems, and AI teams to ensure software fully enables our client's processor roadmap.

Key Responsibilities

Compiler Architecture & Development

  • Design, develop, and optimize compiler infrastructure for custom processor and accelerator architectures
  • Build compiler backends, optimization passes, code generation workflows, and execution pipelines
  • Improve instruction scheduling, register allocation, memory access efficiency, and execution performance
  • Develop graph-level and operator-level optimization strategies for AI workloads
  • Support compiler enablement for emerging compute architectures

Runtime & Performance Optimization

  • Develop runtime systems and execution frameworks for optimized inference and compute performance
  • Profile workloads and identify bottlenecks impacting latency, throughput, memory bandwidth, and power efficiency
  • Debug performance issues across simulation, emulation, and silicon environments
  • Build internal benchmarking and performance analysis tools

Hardware / Software Co-Design

  • Partner with architecture and hardware teams on next-generation processor development
  • Analyze workload behavior to influence architecture decisions
  • Help drive tradeoff analysis involving performance, memory efficiency, latency, and power consumption
  • Collaborate with silicon teams during bring-up and optimization cycles

AI / ML Workload Enablement

  • Optimize execution of machine learning and inference workloads on custom accelerators
  • Support deployment of workloads involving:
    • computer vision
    • edge AI
    • radar
    • sensing
    • autonomous systems
    • signal processing
  • Collaborate with AI software teams on model optimization and framework integration

Required Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related discipline
  • 7+ years of relevant experience (Senior level)
  • 10+ years of relevant experience (Staff level)
  • Strong C/C++ programming expertise
  • Deep experience with compiler development and optimization
  • Hands-on expertise with one or more of:
    • LLVM
    • MLIR
    • TVM
    • XLA
    • GCC
    • custom compiler frameworks
  • Experience in:
    • compiler backend development
    • code generation
    • optimization passes
    • runtime systems
    • performance profiling
  • Strong understanding of:
    • computer architecture
    • memory systems
    • instruction scheduling
    • register allocation
    • parallel execution
    • low-level performance optimization

Preferred Qualifications

  • Semiconductor industry experience
  • Experience with AI accelerators, NPUs, DSPs, GPUs, or custom compute architectures
  • Knowledge of RISC-V architectures
  • Experience with:
    • graph optimization
    • quantization
    • inference optimization
    • edge AI deployment
    • embedded systems
    • HW/SW co-design
  • Familiarity with autonomous systems, radar, signal processing, or computer vision workloads
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary