Lead Software Engineer - Python/Go & AI/ML
Listed on 2026-09-03
-
Software Development
AI Engineer (Applied/Software)
At JPMorgan
Chase, we are building the infrastructure that powers the next generation of enterprise AI — and we need talented engineers who are passionate about LLM inference to help us do it. This is your opportunity to work at the intersection of cutting-edge machine learning and large-scale production systems, directly contributing to how one of the world's largest financial institutions deploys and optimizes AI at scale.
As a Lead Software Engineer at JPMorgan
Chase within the AI/ML Data Platform team, you will be a key technical contributor on LLM inference performance — supporting optimization strategy, benchmarking, and efficiency will collaborate closely with senior engineers and engineering leadership to help shape how our platform evolves, ensuring every model we serve is fast, cost-efficient, and production-ready. This is a high-impact individual contributor role where your technical contributions will have direct, measurable influence on the firm's AI capabilities.
- Execute systematic benchmarking and performance characterization across production LLM workloads, establishing reproducible baselines, identifying regressions, and quantifying the impact of configuration changes before they reach production
- Design and run quantization experiments — FP8, INT8/INT4 (GPTQ/AWQ), and next-generation precision formats — measuring accuracy delta, throughput improvement, memory reduction, and cost-per-token impact
- Support speculative decoding strategy across the model portfolio, including draft model, n-gram, and multi-token prediction approaches, contributing to acceptance rate measurement and per-workload configuration recommendations
- Build and maintain GPU efficiency metrics covering utilization, memory headroom, cost per 1K tokens, and waste identification — providing engineering teams with a data-driven view of platform efficiency
- Benchmark the platform against external providers and published industry numbers, identifying gaps and contributing to improvement initiatives
- Participate in inference engine upgrade evaluations, including new scheduler architectures, async tensor parallelism, disaggregated prefill/decode, and advanced speculative decoding, supporting systematic validation before production promotion
- Contribute to GPU chaos engineering efforts, including induced failure scenarios, hardware diagnostic monitoring, and detection and recovery measurement
- Leverage enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards
- Apply knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation
- Formal training or certification on software engineering concepts and advanced applied experience – preferably Go / Python
- Hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D, or equivalent production serving engines
- Strong understanding of GPU memory architecture, including KV cache sizing and dynamics, memory-bandwidth versus compute bottlenecks, and the practical implications of quantization at inference time
- Experience with quantization techniques and their real-world tradeoffs at scale
- Familiarity with speculative decoding and the variables that drive acceptance rates in production workloads
- Rigorous benchmarking skills using GuideLLM, custom harnesses, or equivalent tooling, with the ability to support every performance claim with data
- Experience operating in cloud GPU infrastructure at scale (AWS, Kubernetes-based managed inference services)
- Ability to communicate technical trade-offs clearly to engineering peers and senior stakeholders
- Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, testing, troubleshooting, or documentation) with demonstrated ability to critically evaluate and validate AI-generated outputs
- Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations
- Experience with disaggregated prefill/decode serving architectures
- Familiarity with GPU hardware diagnostics tools such as DCGM, NVML, or XID event tracking
- Experience with ML observability and production monitoring for inference workloads
- Awareness of the LLM inference competitive landscape with a track record of applying industry benchmarks to drive platform improvements
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: