Lead Engineer - Infrastructure & Cloud Engineering
Listed on 2026-08-27
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Infrastructure
About Us
Core
42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core
42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond.
The Lead Engineer - Infrastructure & Cloud Engineering provides technical leadership for the design, implementation, operation, and continuous evolution of large-scale private cloud, virtualization, container, and observability platforms across Core
42 infrastructure environments.
The role combines deep hands-on technical expertise with technical leadership responsibilities, driving platform strategy, architecture alignment, engineering standards, operational excellence, and adoption of emerging technologies across Infrastructure Engineering teams.
Responsibilities- Lead the technical design, architecture, implementation, and lifecycle management of large-scale private cloud, virtualization, container, and observability platforms based on Open Stack, Red Hat Open Shift, and related infrastructure technologies.
- Define and drive engineering standards, design principles, operational best practices, and technical governance frameworks across infrastructure and cloud platforms.
- Provide technical leadership and mentorship to engineering teams, supporting technical decision-making, problem-solving, knowledge sharing, and continuous skill development.
- Lead the strategy, design, and adoption of observability platforms and services, including metrics, logs, traces, dashboards, alerting, SLOs/SLIs, and service health visibility capabilities.
- Drive the adoption of AI-assisted operations and Agentic AI technologies to improve operational efficiency, incident management, root cause analysis, platform reliability, and service automation.
- Lead the evaluation, integration, and adoption of new technologies, products, and architectural approaches in collaboration with Architecture, Product Engineering, SRE, Security, and Operations teams.
- Own technical oversight of complex platform upgrades, migrations, production changes, incident response activities, root cause analysis, and reliability improvement initiatives.
- Drive capacity planning, scalability strategies, performance optimization, resiliency engineering, and operational readiness activities across cloud and infrastructure platforms.
- Ensure observability, automation, operational tooling, and platform services are effectively integrated with ITSM, service management, incident management, and operational governance processes.
- Partner with security, compliance, and risk management teams to ensure infrastructure platforms meet cybersecurity, governance, regulatory, and data sovereignty requirements.
- Define and champion automation strategies using Infrastructure as Code, Git Ops, CI/CD, platform engineering, and infrastructure lifecycle management best practices.
- Lead the development and maintenance of technical standards, architecture documentation, operational procedures, design guides, and knowledge management repositories.
- Serve as a senior technical escalation point for critical production incidents and participate in on-call duty rotations, leading complex troubleshooting, incident response coordination, root cause analysis, and service recovery activities.
- Collaborate with leadership teams on infrastructure strategy, roadmap planning, investment decisions, and long-term platform evolution.
- Bachelor's or Master's degree in Computer Science, Engineering, Software Engineering, or a related technology discipline; or equivalent practical experience.
- 8+ years of experience designing, implementing, operating, troubleshooting, and leading large-scale private cloud, virtualization, infrastructure, observability, or platform engineering environments.
- Deep hands-on expertise and technical leadership experience with at least one major cloud or virtualization platform, such as Open Stack, Proxmox, Red Hat Open Shift, or equivalent technologies.
- Experience driving the adoption of AI-assisted operations, workflow automation, and integration of observability and ITSM platforms to improve operational efficiency, incident response, and service reliability.
- Strong expertise in compute technologies, including x86 hardware, KVM, Linux operating systems, hypervisors, firmware, server lifecycle management, and orchestration services.
- Strong hands-on experience with at least one major cloud or virtualization platform, such as Open Stack, Proxmox, Red Hat Open Shift, or equivalent technologies.
- Hands-on experience with AI-assisted operations, workflow automation, and integration of observability and ITSM platforms to improve operational…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).