Principal Software Engineer - LLM Optimization
Listed on 2026-09-30
-
Software Development
AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
At JPMorgan
Chase, we are building the infrastructure that powers the next generation of enterprise AI — and we need the best minds in LLM inference to help us do it. This is your opportunity to work at the intersection of cutting-edge machine learning and large-scale production systems, directly influencing how one of the world's largest financial institutions deploys and optimizes AI at scale.
As a Principal Software Engineer at JPMorgan
Chase within the AI/ML Data Platform team, you will serve as the firm's deepest technical voice on LLM inference performance — owning optimization strategy, benchmarking rigor, and efficiency will work directly with senior engineering leadership to shape how our platform evolves, ensuring every model we serve is fast, cost-efficient, and production-ready. This is a high-visibility individual contributor role where your technical decisions will have direct, measurable impact on the firm's AI capabilities
- Own systematic benchmarking and performance characterization across all production LLM workloads. Establish reproducible baselines, catch regressions early, and quantify the impact of every configuration change before it touches production
- Design and execute quantization experiments — FP8, INT8/INT4 (GPTQ/AWQ), next-generation precision formats on current hardware — measuring accuracy delta, throughput improvement, memory reduction, and cost-per-token impact
- Drive speculative decoding strategy across the model portfolio: draft model, n-gram, and multi-token prediction approaches. Own acceptance rate measurement and per-workload configuration recommendations
- Build and maintain a GPU efficiency scorecard: utilization, memory headroom, cost per 1K tokens, and waste identified — giving leadership a data-driven view of platform efficiency at all times
- Benchmark our platform against external providers and published industry numbers — know what good looks like, and close the gap
- Lead inference engine upgrade evaluations: new scheduler architectures, async tensor parallelism, disaggregated prefill/decode, advanced speculative decoding — systematic validation before production promotion
- Collaborate with the EKS and disaggregated serving teams on KV-cache optimization, prefix caching strategies, and multi-node serving architecture
- Design and run GPU chaos engineering: induced failure scenarios, hardware diagnostic monitoring, detection and recovery measurement
- Architect and govern agentic AI-enabled engineering workflows (using enterprise-authorized tools within the work environment) to improve delivery speed, code quality, and operational outcomes at scale (e.g., AI-driven PR review assistance, test generation/maintenance, release readiness checks, incident triage and root-cause acceleration), while defining guardrails for validation, security, resiliency, and reuse across teams.
- Apply knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation at scale.
- Formal training or certification on software engineering concepts and 7+ years applied experience
- Deep, hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D or equivalent production serving engines
- Strong grasp of GPU memory architecture: KV cache sizing and dynamics, memory-bandwidth vs compute bottlenecks, the practical implications of quantization at inference time
- Experience with quantization techniques and their real-world tradeoffs at scale
- Familiarity with speculative decoding and the variables that drive acceptance rates in production workloads
- Rigorous benchmarking instincts —…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).