×
Register Here to Apply for Jobs or Post Jobs. X

AI Accelerator, Software Engineer- Graph Optimization​/Compilers

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: Ampere
Full Time position
Listed on 2026-09-15
Job specializations:
  • Software Development
    Software Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 159000 - 239000 USD Yearly USD 159000.00 239000.00 YEAR
Job Description & How to Apply Below

Invent the future with us.

Ampere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.

Description

As a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.

Join us at Ampere and work alongside a passionate and growing team -

About

The Role

As a Software Engineer on Ampere’s AI Accelerator team, you will optimize deep learning computational graphs to maximize the performance, efficiency, and scalability of Ampere’s AI accelerator hardware. You will work across the software stack, from model frameworks and inference-serving systems to graph optimization, compiler infrastructure, runtimes, and compute kernels.

What You’ll Achieve
  • Optimize computational graphs for performance, throughput, latency, memory efficiency, and power efficiency on Ampere AI accelerators
  • Enable and optimize models, frameworks, and inference platforms, including PyTorch, Llama.cpp, vLLM, and SGLang
  • Develop graph-level optimizations such as operator fusion, pattern matching, redundancy elimination, constant folding, layout optimization, memory planning, quantization, and accelerator offload
  • Optimize transformer and LLM workloads, including dynamic shapes, attention mechanisms, KV-cache management, and mixed-precision execution
  • Analyze end-to-end performance across frameworks, compilers, runtimes, kernels, and hardware
  • Build profiling, benchmarking, validation, and performance-regression infrastructure
  • Identify bottlenecks using traces, compiler diagnostics, microbenchmarks, and hardware performance data
  • Collaborate with compiler, runtime, kernel, architecture, hardware, and applications teams on hardware/software co-design
  • Contribute to architecture, design reviews, code reviews, documentation, and engineering best practices
About You
  • Bachelor’s degree in Computer Science, Computer Engineering, Mathematics, or a related technical field & 5 years of relevant experience; or a Master’s degree with & 3 years of relevant experience
  • Strong foundations in algorithms, data structures, graph algorithms, computational complexity, and systems programming
  • Proficiency in Python and C/C++, demonstrated through internships, research, coursework, open-source contributions, or personal projects
  • Strong ability to reason about execution dependencies, memory movement, numerical correctness, and hardware execution behavior
  • Experience diagnosing performance issues through profiling, benchmarking, tracing, or hardware-level analysis is a plus
  • Familiarity with deep learning concepts, neural-network architectures, tensor operations, numerical precision, quantization, and memory layout
  • Experience with CUDA, ROCm, OpenCL, SYCL, Triton, GPU programming, NPU programming, or other accelerator architectures is a plus
  • Familiarity with transformer models, LLM inference, attention mechanisms, KV-cache optimization, speculative decoding, mixed-precision execution, or sparsity is a plus
  • Demonstrated exceptional problem-solving ability—IOI medal, ACM ICPC medal, Code forces Grand master, USACO Platinum, or equivalent achievement in research or production engineering is a strong plus
  • Strong analytical and debugging skills, with the ability to investigate ambiguous technical problems and deliver robust solutions
  • Fast learner who can quickly understand new architectures, frameworks, compilers, and workloads
  • Experience using AI-assisted development tools to accelerate implementation, testing, debugging, and code review while maintaining technical ownership and code quality
What We’ll Offer

At Ampere we believe in…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary