AWS DevOps Engineer/AI Platform Engineer; Hybrid Onsite
Listed on 2026-09-05
-
IT/Tech
AI Engineer (Applied/Software), AWS, Cloud Computing: Infrastructure & Operations
AWS Dev Ops Engineer / AI Platform Engineer (Hybrid Onsite)
Location
- Reading, PA
Duration - 6 Months
Interview Type
- Virtual / In-Person
Note
- Hybrid 2-3 days/wk at client office
Description
- Requirements
Experience:
8+ years in Platform Engineering, Dev Ops, or Site Reliability Engineering (SRE).
Cloud Expertise:
Deep proficiency in AWS (IAM, Cloud Watch, Bedrock, Lambda).
Observability Tools:
Proven experience with Dynatrace, Jaeger, or Honeycomb, and distributed tracing standards.
AI/LLM Interest:
Familiarity with the LLM lifecycle, including prompt execution, token usage, and frameworks like Lang Chain or Agent Core.
Automation:
Advanced experience with Terraform and CI/CD pipeline design.
Collaboration:
Experience working in an Agile environment with integrated tools like Microsoft Teams and Confluence.
User this when submitting candidates:
Also please check if the next candidate has some experience with at least 50% of below items:
- Implementation of Agents on Agentcore runtime
- Implementation of Agentic SDLC in Agentcore
- Understanding of Strands or any other Agentic AI framework like Langraph, Langchain or Crew AI
- Implementation of Bedrock Knowledge Base
- Implementation of Knowledge Graph
- Implementation of MCP servers in Agentcore
- Implementation of Agentcore Gateway
- Implementation of Agentcore Identity
- Implementation of Agentic AI Observability
- Implementation of Agentcore Evaluations
- Implementation of AWS Bedrock
- Implementation of AWS Bedrock Inference Profile
- Implementation of AWS Sagemaker
- AWS Services (Cloud) in General
- Terraform
- Deliverables
- Observability:
Assess Cloud Watch, X-Ray, Bedrock logging, Agent Core traces vs. agentic workflow requirements; produce gap analysis, Setup observability in Dynatrace - Design post-deployment validation pipeline for agents & MCP servers (deployment health + tool registration checks)
- Implement distributed tracing & structured logging: LLM decisions, tool selections, sub-agent calls, MCP interactions
- Evaluate Lang Fuse / LiteLLM proxy vs. AWS-native; deliver target-state observability architecture recommendation
- Cost Tracking & TCO:
Extend tagging taxonomy to cover agent runtimes, MCP servers, vector DBs, Bedrock token consumption per namespace - Design cost visibility model: aggregate agent, MCP, vector DB, and Bedrock token costs per team/department
- Build Cloud Watch (or equivalent) dashboards for per-team spend; configure AWS Budgets with alerting thresholds
- Automate cost reports delivered via email / Microsoft Teams; implement anomaly detection rules
- Monitoring & Alerting:
Define P1–P4 alerting rules: deployment failures, runtime errors, tool invocation failures, MCP connectivity issues - Integrate alert notifications to Microsoft Teams channels and email; route by resource ownership tags
- Author runbooks linked to every alert; publish in Confluence for developer self-service resolution
- Evaluate AWS-native vs. third-party monitoring stack; deliver recommendation aligned to observability architecture
Top Skills
- Skill Experience required
Platform Engineering / Dev Ops / SRE8+ years
AWS Cloud Services7+ years
Observability Tools7+ years
AI/LLM Platforms7+ years
Terraform / CI-CD7+ years
Agile Collaboration7+ years
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).