Senior Production Support Engineer - Apigee
Listed on 2026-09-21
-
IT/Tech
SRE/Site Reliability, Systems Engineer, IT Support, AWS
Job Title:
Senior Production Support Engineer - Apigee
Duration:
Full Time / Permanent Position
Location:
Lafayette, LA, Knoxville, TN, Columbia, SC, Birmingham, AL
Work Mode: 5 Days Onsite
Position DescriptionSystemone is seeking a Senior Production Support Engineer to provide advanced operational support for a highly available enterprise API platform running on Kubernetes and Apigee Hybrid within an AWS environment. Candidates must have deep Kubernetes and Apigee experience is required rather than general AWS support alone.
This role is responsible for troubleshooting complex Kubernetes workloads, API gateway and runtime issues, routing and backend connectivity, certificates/TLS, DNS, HTTP, and other platform related production problems. The engineer will lead technical triage, coordinate resolution across application, infrastructure, network, cloud, and Level 3 engineering teams, and support releases, platform changes, failover activities, and production recovery.
The position is primarily focused on production support and operational engineering, with targeted scripting and automation rather than application feature development.
Your future duties and responsibilities- Provide senior level production support for Kubernetes and Apigee Hybrid environments.
- Troubleshoot Kubernetes pods, deployments, replica sets, services, configurations, name spaces, and runtime issues.
- Use kubectl to investigate workload health, application failures, and configuration problems.
- Troubleshoot Apigee API Gateway, including API proxies, routing, target endpoints, runtime components, and backend integrations.
- Investigate certificates, TLS, DNS, HTTP, and application/network connectivity issues.
- Use curl, ping, trace route, and related utilities to diagnose API and network problems.
- Use AWS, Linux, Splunk, and Cloud Watch to investigate complex production issues.
- Lead technical triage and coordinate incident resolution across engineering and support teams.
- Communicate technical findings, business impact, risks, mitigation plans, and recommended actions.
- Support planned releases, platform changes, failover activities, and production recovery.
- Maintain technical documentation and contribute to operational, automation, and resiliency improvements.
- Participate in rotating 24x7 primary on call coverage.
- 8+ years of hands on Kubernetes administration and production troubleshooting experience.
- Hands on Apigee API Gateway experience;
Apigee Hybrid experience is strongly preferred. - Advanced Linux troubleshooting skills and strong hands on experience with kubectl.
- Strong knowledge of Kubernetes pods, deployments, replica sets, services, configurations, and name spaces.
- Experience troubleshooting API proxies, routing, target endpoints, certificates, TLS, and backend integrations.
- Working knowledge of TCP/IP, DNS, and HTTP, along with curl, ping, and trace route.
- Experience supporting applications and platforms in an AWS production environment.
- Experience with technical incident management, production troubleshooting, and complex incident triage.
- Ability to use logs, metrics, dashboards, and alerts to isolate production issues.
- Strong communication and stakeholder management skills, particularly during production incidents.
- Ability to work independently on complex investigations and coordinate effectively with Level 3 and engineering teams.
?
Preferred: AWS Associate certification, Splunk, Cloud Watch, Git Lab, CI/CD, and experience supporting API platforms within a regulated enterprise.
Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field.
: #404-IT Pittsburgh
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).