More jobs:
Infrastructure Services Engineer; Hybrid Eligible
Job in
Oak Ridge, Anderson County, Tennessee, 37830, USA
Listed on 2026-08-25
Listing for:
UT-Battelle
Full Time
position Listed on 2026-08-25
Job specializations:
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure, SRE/Site Reliability
Job Description & How to Apply Below
Overview
We are seeking an Infrastructure Services Engineer who will focus on specializing in monitoring and observability. This position resides in the Infrastructure Operations Center (IOC) in the Digital Services Infrastructure & Operations division of the Information Technology Services Directorate, at Oak Ridge National Laboratory (ORNL).
Major Duties/Responsibilities- Design, implement, administer, and maintain enterprise monitoring and observability solutions across on-premises, cloud, and containerized environments.
- Develop and optimize alerts, dashboards, reports, synthetic monitors, and telemetry pipelines to identify degradation early and accelerate incident triage and root-cause analysis.
- Evaluate monitoring coverage, gaps, overlaps, and underused capabilities, and implement tools, integrations, and data sources that improve operational visibility.
- Automate monitoring deployment, configuration, data collection, and remediation using Power Shell, Python, or other scripting and automation tools.
- Evaluate and apply AI-assisted capabilities for anomaly detection, predictive analytics, and operational efficiency.
- Collaborate with infrastructure, network, application, security, and other technical teams to improve system health and observability.
- Support incident and problem management by providing relevant metrics, logs, performance trends, and historical analysis.
- Work with vendors and internal subject matter experts to troubleshoot monitoring agents, collectors, integrations, and platform components.
- Establish and maintain monitoring standards, topology diagrams, technical documentation, runbooks, and team procedures.
- Support patching, backup, upgrade, and lifecycle activities for monitoring platforms and related infrastructure components.
- Continuously improve alert thresholds, dashboards, data quality, automated remediations, and monitoring workflows to reduce noise and repetitive operational work.
- Provide escalated support for monitoring-related issues and participate in an on-call or planned maintenance rotation as required.
- Deliver ORNL’s mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace – in how we treat one another, work together, and measure success.
- BS degree in information technology or a related technical field and 2 years of relevant experience.
- Experience operating, administering, or engineering enterprise monitoring platforms for infrastructure, applications, networks, or cloud environments.
- Experience supporting enterprise Windows and Linux server environments, including performance analysis and troubleshooting.
- Experience developing automated solutions using Power Shell, Python, or similar scripting tools.
- Working knowledge of cloud infrastructure, container platforms, orchestration technologies, virtualization, and virtual-machine lifecycle operations.
- Understanding of networking fundamentals, system performance indicators, telemetry, and diagnostic methodologies.
- Strong analytical and problem-solving skills, including the ability to use operational data to identify issues and recommend improvements.
- Strong written and verbal communication, customer service, collaboration, and technical documentation skills.
- Ability to prioritize responsibilities and balance project work, operational support, and incident response in a fast-paced environment.
- Demonstrated experience automating repetitive work or improving technical and operational processes.
- Experience engineering and administering one or more enterprise-scale monitoring platforms, such as Prometheus, Grafana, Elastic, Solar Winds, or Dynatrace.
- Experience with observability concepts and technologies, including metrics, logs, traces, baselining, synthetic monitoring, and service-level objectives.
- Experience with anomaly detection, predictive analytics, AIOps, or automated remediation.
- Knowledge of automation and infrastructure-as-code frameworks, such as Ansible, Terraform, or Azure Automation.
- Experience using version-control systems to maintain scripts, configurations, dashboards, or infrastructure code.
- Experience with virtualized or clustered compute environments, including performance tuning and lifecycle automation.
- Familiarity with enterprise storage technologies, including direct-attached, SAN, and object storage, and their monitoring requirements.
- Knowledge of enterprise server, storage, network hardware, and platform-level instrumentation.
- Experience with enterprise backup, patching, configuration, or lifecycle management practices.
- Understanding of change management and controlled operational workflows.
- Experience working in regulated, scientific, government, or similarly complex technical environments.
- Motivated self-starter with the ability to work independently and participate creatively in collaborative teams across the laboratory.
- Visa…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×