×
Register Here to Apply for Jobs or Post Jobs. X

Platform Engineer, AI​/ML Infrastructure

Job in Bethesda, Montgomery County, Maryland, 20811, USA
Listing for: Pfizer
Full Time position
Listed on 2026-09-19
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, AWS
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below
** Staf
* *** f Platform Engineer, AI/ML Infrastructure
** Department:

AI Software & Operations
** Role Summary
** The Staff Platform Engineer, AI/ML Infrastructure will provide technical leadership for thecloud platforms, deployment systems, and operational foundations that power enterprise-scale generative AI applications.  This role will define and evolve the infrastructure architecture for AI/ML platforms running across AWS,Kubernetes, serverless, and containerized environments. The engineer will lead platform standards for reliability, scalability, observability, CI/CD, security, and developer enablement, while partnering closely with software engineering, AI engineering, security, and operations teams.  

The ideal candidate combines deep hands-on cloud engineering experience with staff-level technical influence. They are comfortable designing infrastructure patterns, writing infrastructure-as-code,improving delivery pipelines, mentoring engineers, and making architectural decisions that raise the operational maturity of AI platforms across multiple teams.  
** Key Responsibilities
** Define and drive the technical strategy for AI/ML platform infrastructure supporting generative AI applications, LLM integrations, model routing, and enterprise AI services.  Architect, build, and operate scalable cloud platforms using AWS services such as EKS, ECSFargate, Lambda, DynamoDB, S3, Open Search, Secrets Manager, Cloud Watch, ALB, and MWAA.  Establish reusable infrastructure patterns using Cloud Formation, Helm, and Terraform to support reliable multi-environment and multi-region deployments.  

Lead CI/CD architecture using Git Hub Actions, reusable workflows, OIDC-based AWS authentication, automated quality gates, deployment promotion, and environment approvals.  Design and improve observability across AI platforms, including Cloud Watch dashboards, logs,alarms, Prometheus/Grafana, Open Search, Langfuse, and LLM-specific operational metrics.  Build platform capabilities for GenAI workloads, including model availability monitoring.  Partner with software engineering teams to improve deployment reliability, rollback strategies,health checks, autoscaling, load testing, and runtime performance.  

Define and enforce security and compliance practices for infrastructure, including IAM permission boundaries, Secrets Manager usage, secret scanning, audit logging, tagging standards, andchange-management controls.  Provide technical leadership for cost optimization, capacity planning, environment standardization,and operational resilience across development, test, production, and sandbox environments.  Mentor engineers, review architecture and infrastructure designs, and influence platform engineering practices across teams.
** Basic Qualifications
** Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field, or equivalent practical experience.  7+ years of experience in Dev Ops, platform engineering, cloud infrastructure, site reliabilityengineering, or software engineering roles.  Strong hands-on experience with AWS/Azure/GCP infrastructure and services, including container,serverless, networking, storage, observability, and security services.  Experience designing and operating production systems on Kubernetes, ECS/Fargate, or comparable container orchestration platforms.  

Proficiency with infrastructure-as-code, especially Cloud Formation, Terraform, Helm, or similar tooling.  Strong CI/CD experience with Git Hub Actions or similar platforms, including reusable workflows,automated testing, deployment gates, and cloud authentication.  Experience building and operating observability solutions using Cloud Watch, Prometheus/Grafana,Open Search, or similar tools.  Strong…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary