Principal Technical Account Manager, AWS Enterprise , NAMER-Sp
Listed on 2026-10-03
-
IT/Tech
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Cloud Computing: Infrastructure & Operations, Systems Engineer
Principal Ai/Ml Hpc Specialist
Amazon Web Services (AWS) is seeking an experienced Principal AI/ML HPC Specialist to join our Technical Account Manager (TAM) team. You'll be at the forefront of solving complex AI HPC implementation challenges, guiding NAMER Research labs to enterprise customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem-solving, you'll help organizations unlock the full potential of artificial intelligence and machine learning technologies — from distributed model training on GPU clusters to production-grade inference at scale.
Key job responsibilities:
- Deliver Strategic Technical Engagements
- Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS Parallel Cluster, the latest GPU-accelerated computing (i.e., P6/P6e, G7/G7e instances), AWS Trainium-based training (Trn3 Ultra Servers), and multi-node NCCL communication tuning over EFA's SRD protocol. - Architect and Validate Innovative Solutions
- Design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, Deep Speed, Megatron-LM), Sage Maker Hyper Pod for managed GPU clusters with automated health checks and node replacement, high-performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance hitect solutions using P6e Ultra Servers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference. - Enable Customer Success
- Support customers in implementing business-critical HPC capabilities including the development of large language model (LLM) (Llama, GPT-class models), physics-informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls for GPU/Trainium-intensive workloads. - Enable Business Critical Outcomes
- Partner with service teams to enhance model training throughput, optimize NCCL collective communications, improve GPU/Trainium utilization across multi-node Ultra Clusters, and drive operational efficiency through proactive monitoring, automated failure recovery (Hyper Pod health checks), and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap, share reference architecture performance and benchmarks with broader TAM and Technical communities. - Serve as Trusted Advisor and Advocate — Develop and nurture technical partnerships with enterprise stakeholders, serving as the trusted advisor for AI/ML infrastructure decisions spanning compute, networking (Elastic Fabric Adapter with SRD), storage, orchestration, and the HPC-to-AI convergence journey.
A day in the life
Your day will be dynamic and impactful involving deep technical consultations on distributed training architectures, strategic solution design for GPU and Trainium cluster deployments, and collaborative problem-solving across multi-node ML environments. You'll engage with technical leaders, architect innovative AI/ML implementations — from Slurm-managed PCS clusters and Sage Maker Hyper Pod to PyTorch FSDP/Deep Speed training jobs and Neuron SDK compilation workflows — and provide expert guidance that bridges machine learning infrastructure with business objectives.
You will partner with TAMs, SAs, and service teams to provide…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).