HCI Sr. Compute Engineer (Red Hat OpenShift
Listed on 2026-08-08
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Location: Town of Poland
Stefanini Group is seeking a Senior Compute Engineer specialised in Red Hat Open Shift to strengthen our Compute Operations team and provide Level 3 (L3) expert support for enterprise customers running critical workloads on container and virtualization platforms. This role is a key technical position focused on day-to-day operations, stability, and continuous improvement of Open Shift-based platforms. The engineer will act as the highest escalation point for complex incidents and problems, support platform lifecycle activities (upgrades, patching, performance tuning), and contribute to platform modernization initiatives – including VMware-to-Open Shift virtualization transformation programs.
The ideal candidate combines strong troubleshooting skills, deep infrastructure understanding, and hands‑on Open Shift expertise, with the ability to work in a structured operational environment (ITIL/managed services), while also supporting automation and standardisation.
- Act as the L3 escalation point for complex technical issues related to:
- Red Hat Open Shift clusters (control plane, worker nodes, networking, storage, authentication)
- Open Shift Virtualization (Kube Virt) and VM-based workloads hosted on Open Shift
- Linux OS level issues impacting cluster stability or workloads.
- Own and drive resolution of:
- Major Incidents (P1/P2) with deep technical investigation and rapid recovery focus
- Recurring incidents through Problem Management (root cause analysis and permanent fixes).
- Lead deep troubleshooting activities:
- cluster degradation, node failures, API instability, etcd performance issues
- networking issues (ingress, routes, DNS, CNI, service connectivity)
- storage issues (persistent volumes, performance bottlenecks, CSI failures)
- workload failures (pods, operators, deployments, stateful applications).
- Provide clear technical updates during incidents, including impact assessment, recovery plan / workaround, risks and next steps.
- Plan and execute Open Shift lifecycle activities such as version upgrades (cluster upgrades and operator upgrades), patching and security hardening and certificate management and renewal processes.
- Validate platform readiness before changes: capacity, compatibility, performance, known issues.
- Maintain high availability and resilience: backup/restore strategy support (including etcd backup practices), disaster recovery readiness and operational runbooks.
- Ensure operational compliance with defined maintenance windows and change governance.
- Support enterprise modernization initiatives involving migration from traditional virtualization platforms (VMware) to Open Shift Virtualization
- Contribute to:
- migration approach definition and technical design support
- workload onboarding, validation, and stabilization on Open Shift
- performance tuning and operational model definition for VM-based workloads on Open Shift.
- Ensure production‑grade operational readiness: monitoring, alerting, backup, patching and support model aligned with managed services standards.
- Develop and maintain operational documentation, including troubleshooting guides, standard operating procedures (SOPs), build standards and reference architectures, operational runbooks for recurring tasks.
- Support automation initiatives using tools such as:
Ansible / Automation Platform (preferred), Git Ops practices (ArgoCD) where applicable and scripting (Bash / Python) to reduce manual operations. - Proactively identify improvements to increase platform stability, recovery speed (MTT), repeatability and reduction of human error.
- Support and improve observability across the platform, including:
- Open Shift monitoring stack (Prometheus / Alert manager / Grafana)
- log management (e.g., EFK / Loki or enterprise logging platforms).
- Troubleshoot performance issues related to compute resource constraints, scheduling and resource requests/limits and cluster scaling and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).