Senior Engineer - Infrastructure and Cloud Engineering
Listed on 2026-08-25
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, IT Infrastructure
Senior Engineer - Infrastructure and Cloud Engineering About Us
Core
42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core
42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond
The Senior Engineer – Infrastructure & Cloud Engineering contributes to the design, implementation, operation, and continuous improvement of large-scale private cloud, virtualization, and observability platforms across Core
42 infrastructure environments.
The role is hands-on and focused on deploying, operating, troubleshooting, automating, and optimizing virtualization and observability platforms capabilities across production and non-production environments.
Responsibilities- Design, implement, and operate observability platforms and services, including metrics, logs, traces, dashboards, alerting, and service health visibility using technologies such as Prometheus, Grafana, Open Telemetry, ELK/Open Search, and related platforms.
- Contribute to the design, implementation, and operation of large-scale private cloud, virtualization, and container platforms based on Open Stack, Red Hat Open Shift, and related infrastructure technologies.
- Develop, integrate, and maintain AI-powered operational agents that leverage observability, monitoring, logging, ITSM, and platform telemetry systems to automate incident detection, root cause analysis, operational workflows, and approved remediation activities in accordance with established operational processes and governance controls.
- Collaborate with architecture, product, platform engineering, SRE, security, and operations teams on technology evaluation, integration, solution design, and operational readiness.
- Support capacity planning, performance optimization, production upgrades, migrations, incident response, root cause analysis, and continuous improvement of platform reliability and operational resilience.
- Integrate observability and platform management capabilities with ITSM, incident, change, and problem management processes and tools such as Jira, Service Now, Pager Duty, Opsgenie, or similar platforms.
- Collaborate with security teams to ensure compute, virtualization, cloud, observability, and management platforms are secure, hardened, and aligned with cybersecurity, compliance, and data governance requirements.
- Create and maintain technical documentation, including implementation guides, operational procedures, dashboards, runbooks, diagrams, standards, and knowledge base articles.
- Prepare and deliver technical knowledge transfer sessions for operational teams, SREs, and engineering stakeholders.
- Participate in on-call rotations and provide technical escalation support for critical production incidents, major service disruptions, and platform emergencies, ensuring timely restoration of services and effective root cause resolution.
- Work with process and operations teams to improve support workflows, service onboarding, operational procedures, and collaboration efficiency.
- Bachelor’s or Master’s degree in Computer Science, Engineering, Software Engineering, or a related technology discipline; or equivalent practical experience.
- 5+ years of hands-on experience designing, implementing, operating, troubleshooting, and managing private cloud, virtualization, infrastructure, observability, or platform engineering environments.
- Strong hands-on experience with at least one major cloud or virtualization platform, such as Open Stack, Proxmox, Red Hat Open Shift, or equivalent technologies.
- Hands-on experience with AI-assisted operations, workflow automation, and integration of observability and ITSM platforms to improve operational efficiency, incident response, and service reliability.
- Strong hands-on experience with compute technologies, including x86 hardware, KVM, Linux operating systems, hypervisors, firmware, server lifecycle management, and orchestration services.
- Expert-level Linux administration skills, including troubleshooting, performance analysis, system tuning, patching, and operational support of Linux-based infrastructure environments.
- Strong understanding of hardware architecture and components, including x86/ARM, NUMA, memory channels, NICs, GPU/AI accelerators, firmware, and large-scale server platforms.
- Good understanding of data center networking concepts, including OSI model, TCP/IP, routing, firewalls, load balancing, VLAN/VXLAN, DNS, DHCP, and related technologies.
- Good understanding of storage types, architectures, and protocols, including object, block, file storage, and storage integration with cloud platforms.
- Hands-on experience with…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).