×
Register Here to Apply for Jobs or Post Jobs. X

AI Infrastructure Engineer

Job in Boston, Suffolk County, Massachusetts, 02298, USA
Listing for: Netpreme
Full Time position
Listed on 2026-09-07
Job specializations:
  • Software Development
    Backend Developer, AI Engineer (Applied/Software), Software Engineer, DevOps
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below

About the Role

We're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.

Essential Duties & Responsibilities
  • Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines.
  • Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems.
  • Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency scaling, GPU utilization, and memory utilization.
  • Analyze model architecture and its serving implications, including attention, KV cache, MoE, long context, and speculative decoding.
  • Tune vLLM and SGLang configurations such as continuous batching, max batched tokens, chunked prefill, prefix caching, KV-cache precision/capacity, speculative decoding, CUDA Graphs, and P/D disaggregation.
  • Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime.
  • Compare deployment configurations and identify production operating points balancing latency, throughput, capacity, and stability.
  • Work with model/system engineers to bring newly released models into production efficiently.
  • Collaborate closely with our hardware/systems team (direct access to CTO-level technical leadership on a small team) to translate performance requirements into backend architecture decisions.
  • Contribute to defining next-generation benchmarks and service requirements as workloads evolve - multi-turn coding, agentic pipelines, RAG, and other long-context use cases.
Qualifications
  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering.
  • Must have: hands-on experience deploying and performance-tuning vLLM and/or SGLang.
  • Strong understanding of LLM inference fundamentals, including prefill vs. decode, batching, KV cache, latency/throughput trade-offs, and distributed GPU execution.
  • Strong Python engineering skills.
  • Working knowledge of inference-serving concepts: continuous batching, KV cache handling, quantization, and serving SLAs.
  • Clear written and verbal communication skills to work effectively with a small, fully distributed team.
Preferred Qualifications (optional)
  • Contributions to vLLM, SGLang, Flash Infer, TensorRT-LLM, etc.
  • Experience with MoE / long-context model deployment.
  • Experience with speculative decoding, prefix caching, P/D disaggregation, attention/KV optimization.
  • Experience with Nsight Systems / PyTorch Profiler.
  • Familiarity with Kubernetes / production GPU serving.
  • Previous startup experience.
Compensation & Benefits
  • Competitive salary with performance-based bonus and early-stage equity grant
  • 100% employer-paid Health, Dental, and Vision coverage for you and your dependents
  • 401(k) match with immediate vesting, and access to financial advisors to help you reach your financial goals
  • 100% employer-paid Life, Disability, and AD&D insurance, plus a fitness stipend and wellness & mental health perks
  • Generous PTO: 20 vacation days, 15 company holidays (including 3 floating days of your choosing)
  • Daily lunch stipend
  • Enterprise-level Claude & ChatGPT access with a generous token budget
  • Well-equipped, sunny offices in Santa Clara, CA & Cambridge, MA with on-site parking and EV charging; on-site fitness center in Santa Clara; gym discounts near our Cambridge office
  • Visa sponsorship and relocation assistance to one of our office hubs
  • A collaborative, continuous-learning environment with smart, dedicated colleagues building the next generation of high-performance computing architecture
The Opportunity
  • Impact: Humanity stands at the dawn of a new industrial revolution driven by AI - one with the potential to redefine how we live on this planet. We are tackling a fundamental challenge at the infrastructure layer: unlocking greater AI capability while dramatically improving efficiency. The work we do here compounds across state-of-the-art AI models, systems, and real-world applications.
  • Timing: Breakthrough technology matters most when it meets the right time. Joining now means real ownership of the company and meaningful influence over product direction and execution. In this early-stage environment, your ideas shape the trajectory of the technology - not just its implementation. You’ll work from first principles, move quickly from insight to execution, and see your contributions directly reflected in what we build.
  • Culture: You’ll work alongside a group of people who care deeply about rigor, clarity, and impact. We value thoughtful disagreement, fast learning, and intellectual fearlessness. This is a place where strong ideas shine, curiosity is encouraged, and growth is a daily practice - not a future promise.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary