×
Register Here to Apply for Jobs or Post Jobs. X

Principal Technical Account Manager, AWS Enterprise , NAMER-Sp

Job in Dallas, Dallas County, Texas, 75219, USA
Listing for: Amazon
Full Time position
Listed on 2026-09-13
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software), AWS, Machine Learning/ ML Engineer
Job Description & How to Apply Below
Position: Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp
Description

As part of the AWS Applied AI Solutions organization, we have a vision to provide business applications, leveraging Amazon's unique experience and expertise, that are used by millions of companies worldwide to manage day-to-day operations. We will accomplish this by accelerating our customers' businesses through delivery of intuitive and differentiated technology solutions that solve enduring business challenges. We blend vision with curiosity and Amazon's real-world experience to build opinionated, turnkey solutions.

Where customers prefer to buy over build, we become their trusted partner with solutions that are no-brainers to buy and easy to use.

Are you ready to transform how businesses leverage artificial intelligence and machine learning at scale? Join our team and become a strategic partner in delivering Amazon AI/ML solutions that empower global enterprises to innovate, optimize, and achieve unprecedented operational excellence.

Amazon Web Services (AWS) is seeking an experienced Principal AI/ML HPC Specialist to join our Technical Account Manager (TAM) team.

You'll be at the forefront of solving complex AI HPC implementation challenges, guiding NAMER Resarch labs to enterprise customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem-solving, you'll help organizations unlock the full potential of artificial intelligence and machine learning technologies - from distributed model training on GPU clusters to production-grade inference at scale.

AWS Support includes experts from across AWS who help our customers design, build, operate, and secure their cloud environments. Customers innovate with AWS Professional Services, upskill with AWS Training and Certification, optimize with AWS Support and Managed Services, and meet objectives with AWS Security Assurance Services. Our expertise and emerging technologies include AWS Partners, AWS Sovereign Cloud, AWS International Product, and AI/ML-native solutions.

You'll join a diverse team of technical experts in dozens of countries who help customers achieve more with the AWS cloud.

Key job responsibilities

Deliver Strategic Technical Engagements - Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads, including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS Parallel Cluster, the latest GPU-accelerated computing (i.e. P6/P6e , G7/G7e instances), AWS Trainium-based training (Trn3 Ultra Servers), and multi-node NCCL communication tuning over EFA's SRD protocol.

Architect and Validate Innovative Solutions - Design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, Deep Speed, Megatron-LM), Sage Maker Hyper Pod for managed GPU clusters with automated health checks and node replacement, high-performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance at scale..

Architect solutions using P6e Ultra Servers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference.

Enable Customer Success - Support customers in implementing business-critical HPC capabilities, including the development of large language model (LLM)  (Llama, GPT-class models), physics-informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls for GPU/Trainium-intensive workloads.

Enable Business Critical Outcomes - Partner with  with service teams to enhance model training throughput, optimize NCCL collective communications, improve GPU/Trainium utilization across multi-node Ultra Clusters, and drive operational efficiency through proactive monitoring, automated failure recovery (Hyper Pod health checks), and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR, share refrerence architecture, performance , and benchmarks with broader TAM and Technical communities

Serve as Trusted Advisor and Advocate - Develop and nurture technical partnerships with enterprise stakeholders, serving as the trusted advisor for AI/ML…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary