More jobs:
AWS Platform Support Lead
Job in
Rahway, Union County, New Jersey, 07065, USA
Listed on 2026-09-29
Listing for:
CYNET SYSTEMS
Full Time
position Listed on 2026-09-29
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, AWS, SRE/Site Reliability, Systems Administrator
Job Description & How to Apply Below
Our client is seeking an experienced AWS Platform Support Lead to oversee cloud platform operations, production support, incident management, and infrastructure reliability across AWS environments. The successful candidate will lead a team of support engineers, drive operational excellence, ensure platform stability and security, and act as the primary escalation point for critical incidents. This role requires strong expertise in AWS cloud services, infrastructure operations, automation, monitoring, and ITIL-based service management processes.
Key Responsibilities:
Lead and mentor a team of AWS Platform Support Engineers. Provide technical guidance for cloud operations, support activities, and platform improvements. Define support standards, operational procedures, and best practices. Serve as the technical escalation point for complex production issues. Manage and support AWS infrastructure across multiple environments. Ensure high availability, scalability, performance, and reliability of cloud platforms. Oversee incident, problem, change, and release management activities.
Drive major incident management and coordinate recovery efforts. Perform Root Cause Analysis (RCA) and implement preventive actions. Establish proactive monitoring and alerting mechanisms. Define and track platform SLAs, KPIs, and operational metrics. Analyze trends and recommend improvements for platform stability. Ensure capacity planning and performance optimization. Drive automation of operational activities using Infrastructure as Code (IaC). Support CI/CD processes and release deployments.
Implement operational automation using Terraform, Cloud Formation, Ansible, Python, and Shell Scripting. Ensure cloud environments adhere to security policies and compliance requirements. Review IAM configurations, access controls, and overall security posture. Collaborate with security teams to remediate vulnerabilities and audit findings. Partner with application, Dev Ops, security, network, and business teams. Provide regular operational reports and service review updates. Lead customer and stakeholder communications during major incidents and service reviews.
Required Technical
Skills:
AWS Services. EC2. S3. VPC. IAM. RDS. Lambda. Route 53. ELB / ALB / NLB. Cloud Watch. Cloud Trail. Auto Scaling. SNS / SQS. EKS / ECS. Linux (RHEL, Amazon Linux, Ubuntu). Windows Server. Client . Datadog. Dynatrace. Prometheus. Grafana. ELK Stack. Terraform. Cloud Formation. Ansible. Jenkins. Git Hub. Git Lab. Azure Dev Ops. Python. Shell Scripting. Power Shell. Amazon RDS.
MySQL. PostgreSQL. DynamoDB. Leadership Responsibilities:
Manage and coach a team of cloud support engineers. Conduct performance reviews and development planning. Drive shift governance and support coverage planning. Ensure adherence to operational SLAs and OLAs. Lead major incident bridges and stakeholder communications. Drive continuous service improvement initiatives.
Required Qualifications:
Bachelor's degree in Computer Science, Information Technology, or a related discipline. 8+ years of IT infrastructure and production support experience. 5+ years of hands-on AWS cloud operations experience. Experience leading platform support teams in a production environment. Strong understanding of cloud networking, security, and infrastructure architecture. Experience working within ITIL-based service management environments. Excellent communication, leadership, and stakeholder management skills. Preferred
Certifications:
AWS Certified Solutions Architect Professional. AWS Certified Sys Ops Administrator Associate. AWS Certified Dev Ops Engineer Professional. AWS Certified Security Specialty. ITIL Foundation Certification. Key
Competencies:
AWS Cloud Operations Leadership. Production Support Management. Incident & Problem Management. Platform Reliability Engineering. Site Reliability Engineering (SRE). Team Leadership & Mentoring. Stakeholder Management. Automation & Continuous Improvement. Cloud Security & Governance. Performance Optimization. Nice to Have:
Experience with Kubernetes (EKS) and container platforms. Exposure to Fin Ops and AWS cost optimization practices. Multi-cloud experience (Azure and/or GCP). Knowledge of SRE practices and observability frameworks. Experience supporting global 24x7 operations. Benefits:
Our Benefits Include:
Medical, Dental, and Vision Insurance 401(k) Retirement Plan Health Savings Account (HSA) Disability Insurance (Short-Term…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×