AWS Production Support Engineer
Listed on 2026-09-16
-
IT/Tech
AWS, IT Support, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Job Title:
AWS Production Support Engineer
Location:
Knoxville, Tennessee
Type:
Contract To Hire
Compensation: $75,000.00
Work Model:
Onsite - onsite
Hours:
40.0
Security Clearance:
Not specified
- Provide production support for AWS hosted enterprise applications.
- Monitor applications and infrastructure using Amazon Cloud Watch, AWS Cloud Trail, Splunk, and related tools.
- Investigate alerts through the AWS console, logs, and metrics to identify failure points and business impact.
- Troubleshoot AWS services and Amazon SNS related events.
- Manage incidents, operational tickets, and support queues through Service Now.
- Perform technical triage and coordinate resolution with application developers, cloud teams, and SMEs.
- Escalate complex incidents with complete technical findings and supporting evidence.
- Follow established knowledge articles while independently investigating issues without documented resolutions.
- Support weekday deployments, application releases, and production change activities.
- Participate in LTO 2/LTO 3 resiliency exercises and scheduled weekend LQS activities.
- Identify vulnerabilities and operational risks encountered during production support.
- Develop knowledge of application functionality and take increasing ownership of recurring incidents.
- Ensure incidents and tickets are progressed or resolved in a timely manner.
- Work effectively within the team's 24x7 production support model.
- 2+ years of hands on AWS cloud operations and production support experience.
- Strong ability to navigate the AWS console and troubleshoot application, service, and infrastructure issues.
- Hands on experience with Amazon Cloud Watch and AWS Cloud Trail for monitoring, logs, metrics, and incident investigation.
- Strong production troubleshooting, log analysis, and incident triage skills.
- Experience with Service Now for incident and ticket management.
- Ability to investigate application alerts, transaction failures, AWS service issues, and operational incidents.
- Experience coordinating incident resolution with application development teams, cloud teams, and technical SMEs.
- Ability to work independently when a knowledge article or documented procedure does not fully address an issue.
- Strong communication, ownership, analytical thinking, and problem solving skills.
- Flexibility to support releases, production changes, resiliency activities, and changing support schedules.
- Ability to work effectively within a 24x7 production environment.
System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.
System One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law.
: #404-IT Pittsburgh
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).