Specialist Solutions Architect - AI, Data & AI GTM
Listed on 2026-09-20
-
IT/Tech
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Cloud Computing: Infrastructure & Operations
Generative AI and large-scale machine learning are redefining what's possible — and AWS is at the center of that transformation. We are looking for a Machine Learning Solutions Architect (ML SA) who will serve as the technical authority on model customization and inference to help customers across the AMERICAS unlock the full potential of foundation models, custom training, and production-scale serving on AWS.
Amazon has invested in AI for over two decades. From the recommendation engines that power to the deep learning behind Alexa, Prime Air, Amazon Go, and our supply chain optimization — machine learning is embedded in everything we build. Now, through Amazon Sage Maker AI, Sage Maker Hyper Pod, Amazon Bedrock, and our purpose-built silicon (Trainium, Inferentia), we are enabling customers to fine-tune, train, and deploy models at unprecedented scale and efficiency.
As a Model Customization & Inference Sage Maker ML SA, you will work directly with customers — from startups to enterprises — to design end-to-end ML architectures that span the full lifecycle: data preparation, distributed training, model fine-tuning (LoRA, PEFT, RLHF), inference optimization, and production deployment. You will operate across all 2 layers of the AWS AI/ML stack:
- Infrastructure & Compute — Sage Maker Hyper Pod, GPU-based EC2, EKS/ECS for ML and Gen AI workloads
- ML Platforms — Amazon Sage Maker AI (training jobs, endpoints, pipelines, MLOps)
You will be the bridge between customers and AWS engineering — translating real-world business problems into scalable ML architectures and feeding critical customer signals back to service teams to shape the product roadmap.
Key job responsibilities- Solution Design & Delivery:
Partner with customers' data science and engineering teams to deeply understand their business objectives, then architect solutions that leverage AWS AI/ML services — with emphasis on model customization (fine-tuning, continued pre-training, distillation) and inference optimization (model compilation, quantization, endpoint auto-scaling, multi-model endpoints). - Technical Leadership:
Serve as the go-to SME on model customization and inference patterns across Sage Maker AI and Sage Maker Hyper Pod. Guide field SAs and customers on best practices for training at scale and deploying models with optimal latency, throughput, and cost. - Customer Adoption & Revenue Impact:
Partner with Specialist SAs, Account Teams, Sales, and Business Development to accelerate adoption of Sage Maker AI across the AMERICAS — directly contributing to pipeline generation, opportunity progression, and revenue attainment. - Thought Leadership & Evangelism:
Author technical blogs, whitepapers, reference architectures, and reusable solution artifacts. Deliver presentations at flagship events (AWS re:Invent, AWS Summits, industry conferences) to establish AWS as the leader in model customization and inference. - Voice of the Customer:
Act as the technical liaison between customers and AWS service teams (Sage Maker). Capture and elevate product feature requests, identify gaps, and drive platform improvements grounded in real-world customer needs. - Community Building:
Develop and scale an internal community of ML subject matter experts across the AMERICAS, fostering knowledge sharing on model customization, inference optimization, and emerging ML patterns.
Your day starts with a whiteboard session alongside a financial services customer's ML engineering team, walking them through a distributed fine-tuning architecture — helping them set up Supervised Fine-Tuning (SFT) with LoRA on a Llama model using Sage Maker Training Jobs across a cluster of P5e instances, configuring FSDP for efficient multi-GPU parallelism, and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).