×
Register Here to Apply for Jobs or Post Jobs. X

Senior Inference Engineer, AIConfigurator

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: Jobtailor
Full Time position
Listed on 2026-07-20
Job specializations:
  • Software Development
    Software Engineer, AI Engineer (Applied/Software), Backend Developer, Python
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

Responsibilities

  • Build and evolve AIConfigurator's core optimization engine for LLM serving, including configuration search, SLA‑aware ranking, efficiency and latency estimation, and Pareto frontier analysis.
  • Build production‑quality Python/Rust APIs, CLIs, SDK surfaces, and web workflows that help users generate strong deployment configurations for NVIDIA GPU clusters.
  • Develop configuration generation systems that emit backend‑specific artifacts for Dynamo, Kubernetes, TensorRT‑LLM, vLLM, and SGLang deployments.
  • Collaborate with inference runtime, performance, benchmarking, and product groups to ensure simulated results correspond with actual deployment performance on H100, H200, B200, GB200, and upcoming NVIDIA platforms.
  • Improve model, hardware, and backend support by integrating performance databases, profiling data, support matrices, and validation tools.
  • Drive software quality through maintainable architecture, schema development, tests, documentation, and automation suitable for open‑source and production users.
  • Convert intricate inference ideas like prefill/decode disaggregation, tensor parallelism, pipeline parallelism, expert parallelism, batching, and KV cache behavior into dependable software abstractions.
Requirements
  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related field, or equivalent experience.
  • 10+ years of relevant software engineering experience.
  • Strong Python/Rust engineering skills, including production APIs, CLI tools, packaging, testing, debugging, and maintainable software development.
  • Experience with GPU computing, distributed systems, ML infrastructure, or high‑performance model serving.
  • Understanding of LLM inference concepts such as batching, latency, efficiency, memory constraints, parallelism strategies, and serving SLAs.
  • Experience working with data‑driven performance analysis, benchmarking, simulation, optimization, or managing resource needs.
  • Ability to collaborate across research, runtime, platform, and customer‑facing engineering teams.
  • Strong written and verbal communication skills, with the ability to explain sophisticated technical tradeoffs clearly.
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary