Site Reliability Engineer; SRE C2C LOCALS
Listed on 2026-06-30
-
IT/Tech
Systems Administrator, IT Support, Cloud Computing: Infrastructure & Operations, Unix/Linux
United Airlines. 100% onsite in Arlington Heights.
Need an admin background (Windows, Linux, AWS…) monitoring experience. Dynatrace is really preferred. SRE direction. More on the dev‑ops/ops type role.
11pm – 7:30am CST. If we call their first day Wednesday they will work Wednesday, Thursday, Friday, Saturday, Sunday.
Job DescriptionThis support position in United’s Digital Technology Command Center requires 100% on‑site shift work providing 24/7 operational support via proactive monitoring. As part of the Application Recovery Team ART this position supports monitoring for Digital Channels Production environments, application health, distributed system support in Unix/Linux Windows Virtual VMWare HyperV middleware and AWS Cloud environments.
Responsible for monitoring day‑to‑day application performance and availability—monitoring all aspects of systems performance and analysis of alerts from various tools, and restoration of services.
Familiarity with standard concepts, practices and procedures within IT operational support. Relies on experience and judgment to plan for shift workload and participate in planned change activities as well as incident follow‑ups.
Creatively managing multiple alerts and incidents to ensure high‑impacting scenarios are prioritized appropriately and addressed quickly to reduce system impact. Being able to correlate events from various tools to assess appropriate level of impact.
Responsible for providing day‑to‑day support for application recovery and availability, Digital Channels system administration via proactive monitoring for Windows and UNIX/Linux HPUX AIX Solaris Linux environments.
Understanding of various Microsoft tools is a plus. Additional monitoring and support required for critical processes in cloud virtualization middleware database storage and backup areas. Troubleshooting and familiarity with various middleware components, messaging technology, Web Logic, Web Sphere, Data Power is essential.
Strong knowledge of enterprise monitoring tools such as Dynatrace, Datadog, App Dynamics, Big Panda, SCOM and Logic Monitor is critical. Familiarity with incident ticketing system Service Now is a plus.
Responding to alerts from enterprise monitoring tools, automation and escalation from IT Service Desk. Work with cross‑functional teams throughout IT to isolate and resolve unplanned outages. Ability in writing scripts using Shell, Python etc. to troubleshoot and automate remediation of alerts to enhance system stability and efficiency is a plus.
This includes implementing best practices for error handling, logging and performance optimization to ensure robust and reliable system operations.
Strong working knowledge of ITIL service management is required. Be able to follow change, incident and problem management activities.
Solid verbal and written communication with other technical teams and the business. Able to work well with application support, Dev Ops, Digital Operations Center, server operations, network operations, middleware and database support teams onshore/offshore, being able to respond to user reported issues.
Proactively identify any issues with application functionality abnormalities in digital channel performance, server hardware and software to ensure stability and availability 24x7x365. Ensure enterprise standards and security are maintained and enforced on hardware and software.
Quickly and effectively support incident response process, provide troubleshooting expertise using various monitoring tools during incidents, engage in problem management follow‑up and implement improvements to prevent similar incidents.
Experience with application performance tools, enterprise .com site and mobile channel monitoring intermediate.
Systems administration experience in Unix/Linux and Windows server environment required. Good understanding of virtual technology is also a must.
Minimum 1-2 years of experience in an operational support role is essential. Strong knowledge of ITSM service management best practices and strong written and communication skills is a must.
Ability to independently solve application performance and system problems, be self‑directed and have…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).