Director, Core Infrastructure Engineering
Job in
Cheyenne, Laramie County, Wyoming, 82009, USA
Listed on 2026-07-24
Listing for:
Oracle
Full Time
position Listed on 2026-07-24
Job specializations:
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
* Oracle Cloud Infrastructure (OCI) is building some of the world's largest and most advanced GPU clusters to power the next generation of AI. The Strategic Customer Engineering (SCE) Core Infrastructure team-also known as the AI/ML Forward Deployed Infrastructure Engineering team-provides white-glove engineering and operational support to OCI's most strategic AI/ML infrastructure customers.
As a trusted partner to our customers, we play a critical role in designing, deploying, operating, and optimizing the infrastructure that powers some of the largest and most demanding GPU and AI/ML environments in the world. Our team works closely with customers and internal engineering organizations to ensure exceptional reliability, performance, and scalability for mission-critical AI workloads.
We are seeking an experienced Core Infrastructure Engineering Leader to lead a high-performing engineering team responsible for delivering healthy infrastructure with optimal performance configuration. This leader will drive end-to-end technical customer execution including; POCs, technical troubleshooting, post-sale customer management/support, build automated solutions for provisioning, configuring, and monitoring AI/ML infrastructure to streamline operations and enhance productivity and optimize infrastructure performance.
LI-ES2
** Responsibilities*
* ** Responsibilities*
* Lead, mentor, and develop a team of Core Infrastructure Engineers responsible for designing, implementing, and maintaining the infrastructure that supports our largest GPU/AI/ML customers.
Drive the design, development, testing, validation, and deployment readiness of our automated GPU Cluster deployment tool (like AWS Parallel Cluster, Azure Cycle Cloud) with Slurm and/or Oracle Kubernetes Engine (OKE) to streamline operations and enhance productivity.
Build collaborative relationships with OCI Services team, customer and sales team to deliver reliable, scalable infrastructures. Act as a technical liaison between customers, core engineering teams, and support.
Work with OCI Strategic customers to grow our business in pre/post sales stages in a technical infra expert role.
Take ownership of problems and work to identify solutions. Ability to think through the solution and identify/document potential issues impacting your customers.
Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre-processing techniques.
Troubleshoot infrastructure performance, scalability, and reliability issues and implement solutions to mitigate risks and minimize downtime.
Document infrastructure designs, configurations, and procedures to facilitate knowledge sharing and ensure maintainability.
As a trusted customer advocate, you will help customers/partners understand best practices around advanced GPU solutions, and how to migrate their workloads to the cloud.
Educate customers of all sizes on the value proposition of Oracle Cloud and participate in deep architectural discussions to ensure solutions are designed for successful deployment in the cloud.
*
* Qualifications:
*
* 7+ years of senior software engineering leadership or related experience.
Strong communication and collaboration skills, with the ability to work effectively in cross-functional teams and convey technical concepts to non-technical stakeholders.
Demonstrated leadership and people management skills.
Proven at building and managing distributed/cloud software engineering solutions.
Demonstrated ability to:
+ Develop short, medium, and long-term plans to achieve strategic objectives.
+ Interact across functional areas with senior management or executives to ensure unit objectives are met.
+ Influence thinking and gain acceptance from others in sensitive situations.
BS or MS degree or equivalent experience relevant to functional area.
Experience using tools like Ansible, Terraform, Python, containerization technologies (e.g., Docker, Kubernetes) and orchestration tools
Solid understanding of networking concepts, security principles, and best practices.
Excellent problem-solving skills, with the ability to troubleshoot complex…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×