×
Register Here to Apply for Jobs or Post Jobs. X

HCI Sr. Compute Engineer (Red Hat OpenShift

Job in Town of Poland, Jamestown, Chautauqua County, New York, 14701, USA
Listing for: Stefanini EMEA
Full Time position
Listed on 2026-08-08
Job specializations:
  • IT/Tech
    Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below
Position: HCI Sr. Compute Engineer (Red Hat OpenShift)
Location: Town of Poland

Stefanini Group is seeking a Senior Compute Engineer specialised in Red Hat Open Shift to strengthen our Compute Operations team and provide Level 3 (L3) expert support for enterprise customers running critical workloads on container and virtualization platforms. This role is a key technical position focused on day-to-day operations, stability, and continuous improvement of Open Shift-based platforms. The engineer will act as the highest escalation point for complex incidents and problems, support platform lifecycle activities (upgrades, patching, performance tuning), and contribute to platform modernization initiatives – including VMware-to-Open Shift virtualization transformation programs.

The ideal candidate combines strong troubleshooting skills, deep infrastructure understanding, and hands‑on Open Shift expertise, with the ability to work in a structured operational environment (ITIL/managed services), while also supporting automation and standardisation.

Job Responsibilities Level 3 Operations & Technical Escalation (Core Responsibility)
  • Act as the L3 escalation point for complex technical issues related to:
  • Red Hat Open Shift clusters (control plane, worker nodes, networking, storage, authentication)
  • Open Shift Virtualization (Kube Virt) and VM-based workloads hosted on Open Shift
  • Linux OS level issues impacting cluster stability or workloads.
  • Own and drive resolution of:
  • Major Incidents (P1/P2) with deep technical investigation and rapid recovery focus
  • Recurring incidents through Problem Management (root cause analysis and permanent fixes).
  • Lead deep troubleshooting activities:
  • cluster degradation, node failures, API instability, etcd performance issues
  • networking issues (ingress, routes, DNS, CNI, service connectivity)
  • storage issues (persistent volumes, performance bottlenecks, CSI failures)
  • workload failures (pods, operators, deployments, stateful applications).
  • Provide clear technical updates during incidents, including impact assessment, recovery plan / workaround, risks and next steps.
Platform Lifecycle Management (Upgrades, Patching, Stability)
  • Plan and execute Open Shift lifecycle activities such as version upgrades (cluster upgrades and operator upgrades), patching and security hardening and certificate management and renewal processes.
  • Validate platform readiness before changes: capacity, compatibility, performance, known issues.
  • Maintain high availability and resilience: backup/restore strategy support (including etcd backup practices), disaster recovery readiness and operational runbooks.
  • Ensure operational compliance with defined maintenance windows and change governance.
VMware-to-Open Shift Virtualization Transformation Support
  • Support enterprise modernization initiatives involving migration from traditional virtualization platforms (VMware) to Open Shift Virtualization
  • Contribute to:
    • migration approach definition and technical design support
    • workload onboarding, validation, and stabilization on Open Shift
    • performance tuning and operational model definition for VM-based workloads on Open Shift.
  • Ensure production‑grade operational readiness: monitoring, alerting, backup, patching and support model aligned with managed services standards.
Standardization, Automation & Operational Improvement
  • Develop and maintain operational documentation, including troubleshooting guides, standard operating procedures (SOPs), build standards and reference architectures, operational runbooks for recurring tasks.
  • Support automation initiatives using tools such as:
    Ansible / Automation Platform (preferred), Git Ops practices (ArgoCD) where applicable and scripting (Bash / Python) to reduce manual operations.
  • Proactively identify improvements to increase platform stability, recovery speed (MTT), repeatability and reduction of human error.
Monitoring, Observability & Performance Management
  • Support and improve observability across the platform, including:
    • Open Shift monitoring stack (Prometheus / Alert manager / Grafana)
    • log management (e.g., EFK / Loki or enterprise logging platforms).
  • Troubleshoot performance issues related to compute resource constraints, scheduling and resource requests/limits and cluster scaling and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary