×
Register Here to Apply for Jobs or Post Jobs. X

Director, Core Infrastructure Engineering

Job in Nashville, Davidson County, Tennessee, 37201, USA
Listing for: Hackajob
Full Time position
Listed on 2026-08-15
Job specializations:
  • IT/Tech
    Systems Engineer
Job Description & How to Apply Below

Director Of Core Infrastructure Engineering

OCI is a trusted infrastructure partner to customers running some of the world's largest and most demanding AI and machine-learning workloads. Our teams design, build, deploy, operate, and continuously optimize large-scale GPU environments where reliability, performance, security, and delivery speed are critical.

Raw Metal Cloud (RMC) is an operating platform for the AI-infrastructure lifecycle, from hardware onboarding and inventory through networking, provisioning, validation, rack certification, and customer handoff.

As Director of Core Infrastructure Engineering, you will lead a high-performing engineering organization responsible for evolving RMC and delivering production-ready GPU infrastructure  will own the strategy, architecture, execution, and operational health of critical infrastructure services while partnering across hardware, networking, data-center operations, platform engineering, security, product, and customer-facing organizations.

This role requires a leader who combines deep technical judgment with organizational leadership and disciplined execution. You will develop engineering leaders, establish clear ownership across services, drive architectural alignment, and ensure that teams deliver measurable improvements in infrastructure readiness, fleet health, performance, reliability, security, and customer outcomes.

Key Responsibilities

System Design & Architecture – System Scalability:

  • Implements strategies across multiple teams or groups for the architecture and design of interdependent scalable distributed systems, including the use of distributed state management tools, ensuring organizational and system demands are met.
  • Spearheads code and/or system optimization initiatives for large-scale data processing and high-throughput requirements across multiple areas, driving improvements that support hyper-scale systems.
  • Facilitates collaborations to define system scalability requirements, ensuring the defined requirements meet customer expectations.
  • Oversees the design of interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).
  • Drives the effective use and implementation of data plane platforms for large-scale data operations.

System Design & Architecture – System Reliability Design:

  • Provides strategic oversight for the architecture of fault-tolerant interdependent systems capable of withstanding in-service updates by overseeing implementation across teams of redundancy, replication, and automatic failover mechanisms.
  • Influences and sets direction for designing systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
  • Leads strategic optimization initiatives for handling network unreliability, including directing the design of load-shedding, throttling, and rate-limiting techniques.
  • Holds teams accountable for leveraging formal verification techniques to verify system designs and conduct peer reviews across teams.
  • Drives the design of systems that are durable and adhere to service level objectives (SLOs), developing standards for availability and durability of other computing services across the department.

System Design & Architecture – System Reliability Performance:

  • Drives strategies for defining key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running systems, ensuring alignment with organizational goals.
  • Directs the creation and customization of complex dashboards, telemetry systems, and alerting mechanisms that proactively monitor and ensure optimal system health across teams.

System Design & Architecture – Correctness / Availability:

  • Implements strategies to effectively determine if systems are meeting functional and correctness requirements, and encourages teams to identify improvement opportunities.
  • Provides thought leadership on processes for formally verifying complex features to ensure system design correctness.
  • Oversees the implementation of data replication and synchronization techniques, ensuring data integrity and availability across the organization.

Operational…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary