Site Reliability Engineer - MN, IA, TX
Listed on 2026-08-07
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support
Site Reliability Engineer
In this contingent resource assignment, you will consult on complex initiatives with broad impact and large-scale planning for Software Engineering. This role functions as a senior Site Reliability Engineer supporting enterprise production environments, platform reliability, observability, automation, incident management, and operational excellence initiatives. The successful candidate will join a platform management organization responsible for L2/L3 production support, focusing on improving reliability, implementing automation, and maturing the team's SRE capabilities.
Key Responsibilities
- Provide L2/L3 support for critical enterprise applications and drive reliability improvements across production environments.
- Lead incident, problem, and change management processes, including root cause analysis and corrective action planning.
- Build and maintain observability solutions using platforms like Splunk, App Dynamics, Grafana, and Big Panda.
- Support CI/CD pipelines and enhance release automation using tools such as Jenkins, Artifactory, UDeploy, and Terraform.
- Develop operational automation and self-healing capabilities, primarily using Ansible, to reduce manual effort.
- Support AI/ML, LLM, and Agentic AI-based application environments, partnering with development teams to manage reliability.
- Manage operational risk and performance across on-prem, hybrid, and public cloud infrastructures.
- Support large-scale Java/.NET applications and assist with troubleshooting for Oracle/MSSQL databases.
Required Qualifications
Experience:
8+ years of experience in Software Engineering, Site Reliability Engineering (SRE), Production Support, or Platform Engineering.
Technical
Skills:
- Expertise with observability platforms (e.g., App Dynamics, Splunk, Grafana, Big Panda, Application Insights).
- Experience with CI/CD tools (e.g., Jenkins, Artifactory, UDeploy, Terraform).
- Automation experience using Ansible.
- Experience supporting AI/ML, LLM, and Agentic AI platforms in production.
- Experience with L2 level troubleshooting using Unix commands.
- Familiarity with leading large-scale production support using ITIL practices.
Preferred Qualifications
- Experience in enterprise banking or another highly regulated industry.
- Knowledge of public, hybrid, and on-prem cloud platforms.
- Experience with resilient system design.
- Background in supporting large-scale Java applications.
- Experience with Oracle and MSSQL database support.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).