Senior Manager, Core Infrastructure Engineering- Nashville, TN
Listed on 2026-09-23
-
IT/Tech
Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon for our OCI Network Control Services Automation. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms (load‑shedding, throttling, rate‑limiting).
Provides oversight for KPIs, telemetry, and moderately complex dashboards; directs design of functional/correctness requirements, fault‑injection tests, and replication/synchronization strategies. Ensures proactive incident management, operational readiness, and on‑call coverage; drives encryption/access control practices, remediation plans, and compliance documentation. Oversees development and maintenance of automation/IaC and partners with teams on change‑management plans enabling safe patching, updates, and rollbacks.
Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon for our OCI Network Control Services Automation. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms (load‑shedding, throttling, rate‑limiting).
Provides oversight for KPIs, telemetry, and moderately complex dashboards; directs design of functional/correctness requirements, fault‑injection tests, and replication/synchronization strategies. Ensures proactive incident management, operational readiness, and on‑call coverage; drives encryption/access control practices, remediation plans, and compliance documentation. Oversees development and maintenance of automation/IaC and partners with teams on change‑management plans enabling safe patching, updates, and rollbacks.
This role is based out of Nashville, TN.
ResponsibilitiesKey Responsibilities System Design & Architecture – System Scalability:
- Manages the development and implementation of scalable distributed systems and components across multiple teams, including the effective use of distributed state management tools.
- Oversees code and/or system optimization efforts for large-scale data processing and high-throughput requirements within and across teams to support hyper-scale systems.
- Guides teams to define scalability requirements for owned components and ensures design and implementation requirements are met.
- Manages the use of data plane platforms to effectively handle large-scale data retrieval, storage, and processing.
- Ensures team accurately designs performance and load testing.
- Manages the strategy for building fault‑tolerant components and systems capable of withstanding in‑service updates by guiding the implementation of redundancy, replication, and automatic failover mechanisms.
- Develops design strategies for systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
- Leads implementation and optimization initiatives across teams for approaches to handle network unreliability, including load‑shedding, throttling, and rate‑limiting.
- Guides teams to design components and systems that are durable and adhere to service level objectives (SLOs), setting expectations for availability and durability of other computing…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).