More jobs:
Cloud Platform Engineer; Agentic AI
Job in
Washington, District of Columbia, 20022, USA
Listed on 2026-08-30
Listing for:
Luxoft
Full Time
position Listed on 2026-08-30
Job specializations:
-
IT/Tech
AWS, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Project description
The project is for one of the world's famous science and technology companies in pharmaceutical industry, supporting initiatives in AWS, AI and data engineering, with plans to launch over 20 additional initiatives in the future.
We are seeking a highly skilled Cloud Engineer to lead the infrastructure design, deployment, and operations of the AI agent orchestration platform on AWS. This role is responsible for building and managing a Kubernetes-native, enterprise-grade platform that supports scalable AI agent workloads across development, QA, and production environments.
- 1. AWS Infrastructure & Architecture
- Design, provision, and manage AWS infrastructure using Terraform, aligned with the AWS Well-Architected Framework
- Core services include:
- Amazon EKS- VPC- IAM
- Application Load Balancer (ALB)- Route 53- AWS Certificate Manager (ACM)2. Kubernetes (EKS) Platform Operations - Own and operate EKS clusters end-to-end:
- Managed node group lifecycle management
- Karpenter-based autoscaling
- Cluster add-on lifecycle upgrades- IRSA (IAM Roles for Service Accounts) configuration
- Multi-AZ high availability and resilience - CI/CD & Git Ops
- Build and maintain automated deployment pipelines using:
- Git Hub Actions
- ArgoCD (Git Ops) - Enable multi-environment deployments:
- Dev QA Production - Implement release strategies:
- Blue/Green deployments
- Canary releases - Security & Compliance
- Integrate AWS-native security and governance controls:
- AWS WAF
- Guard Duty
- Security Hub- KMS (encryption)- Secrets Manager
- External Secrets Operator - Enforce policy controls using:
- OPA / Kyverno (admission controllers) - Observability & Monitoring
- Implement and manage observability stack:
- Amazon Managed Prometheus
- Amazon Managed Grafana
- Cloud Watch Container Insights- AWS X-Ray (distributed tracing) - AI/ML Integration
- Leverage AWS AI/ML services to support agent orchestration:
- Amazon Bedrock (model inference, agent APIs)- Sage Maker (model hosting, endpoints)- Comprehend (NLP, PII detection) - Cost Optimization (Fin Ops)
- Implement cost-efficient architecture practices:
- Spot Instances
- Savings Plans
- Karpenter bin-packing strategies
- Scheduled scale-to-zero for non-production environments - Platform & Engineering Collaboration
- Partner with platform and ML teams to:
- Onboard new AI agent workloads
- Integrate MCP servers and execution frameworks
- Support extensibility of the agent ecosystem
- Experience & Certifications
- 4+ years of hands-on AWS experience
- AWS
Certifications:
- Required:
AWS Solutions Architect (Associate or Professional)- Preferred:
Dev Ops Engineer, Security Specialty - Kubernetes & EKS Expertise
- Strong hands-on experience with:
- EKS cluster provisioning and operations
- Managed node groups and Karpenter
- Helm chart management
- Kubernetes RBAC and network policies
Infrastructure as Code (Terraform) - Advanced Terraform capabilities:
- Modular design
- Remote state management (S3 + DynamoDB)- Multi-environment configuration
- Security scanning (Checkov, tfsec) - AWS Services Proficiency
- Deep knowledge of:
- EKS, ECR, ALB, Route 53, ACM- IAM, KMS, Secrets Manager- IAM Identity Center
- Cloud Trail, AWS Config
- Guard Duty, Security Hub, AWS WAF AI/ML Exposure - Practical experience with:
- Amazon Bedrock (model invocation, agent APIs)- Sage Maker (model deployment and endpoints)- Comprehend (NLP and PII detection)
Dev Ops & Identity - Experience with:
- Git Ops tools (ArgoCD or Flux)- CI/CD pipelines for container workloads- OIDC federation:
- Git Hub Actions AWS- EKS OIDC provider integration
Observability & Debugging - Familiarity with:
- Prometheus, Grafana
- Open Telemetry- AWS X-Ray
- Cloud Watch Logs Insights Kubernetes Security - Strong understanding of:
- Pod Security Standards
- Network Policies
- Admission webhooks
- Service account least-privilege principles
- Experience with AI agent frameworks:
- Lang Chain, Claude Agent SDK, or similar - Knowledge of emerging protocols:
- A2A (Agent-to-Agent)- MCP (Model Context Protocol) - Familiarity with:
- Amazon Bedrock Agents, Knowledge Bases, Guardrails - Chaos engineering exposure:
- AWS Fault Injection Service (FIS) - Multi-tenant platform design:
- Namespace isolation
- Self-service provisioning - Programming/debugging skills:
- Python, Go, or Node.js - Fin Ops experience:
- AWS Cost Explorer
- Compute Optimizer
- Tagging governance
- Savings Plan management
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×