×
Register Here to Apply for Jobs or Post Jobs. X

Senior Manager, Core Infrastructure Engineering

Job in Atlanta, Fulton County, Georgia, 30309, USA
Listing for: Oracle
Full Time position
Listed on 2026-08-22
Job specializations:
  • IT/Tech
Job Description & How to Apply Below
** Job Description*
* Manages team delivering scalable distributed systems and components on a 2-4 quarter horizon. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms (load‑shedding, throttling, rate‑limiting). Provides oversight for KPIs, telemetry, and moderately complex dashboards;

directs design of functional/correctness requirements, fault‑injection tests, and replication/synchronization strategies. Ensures proactive incident management, operational readiness, and on‑call coverage; drives encryption/access control practices, remediation plans, and compliance documentation. Oversees development and maintenance of automation/IaC and partners with teams on change‑management plans enabling safe patching, updates, and rollbacks.

** Responsibilities*
* ** Key Responsibilities*
* ** System Design & Architecture - System Scalability:*
* + Manages the development and implementation of scalable distributed systems and components across multiple teams, including the effective use of distributed state management tools.

+ Oversees code and/or system optimization efforts for large-scale data processing and high-throughput requirements within and across teams to support hyper-scale systems.

+ Guides teams to define scalability requirements for owned components and ensures design and implementation requirements are met.

+ Manages the use of data plane platforms to effectively handle large-scale data retrieval, storage, and processing.

+ Ensures team accurately designs performance and load testing.

** System Design & Architecture - System Reliability Design:*
* + Manages the strategy for building fault-tolerant components and systems capable of withstanding in-service updates by guiding the implementation of redundancy, replication, and automatic failover mechanisms.

+ Develops design strategies for systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.

+ Leads implementation and optimization initiatives across teams for approaches to handle network unreliability, including load-shedding, throttling, and rate-limiting.

+ Guides teams to design components and systems that are durable and adhere to service level objectives (SLOs), setting expectations for availability and durability of other computing services within the department.

** System Design & Architecture - System Reliability Performance:*
* + Provides oversight in defining key performance indicators (KPIs) and telemetry to identify gaps or issues in running systems.

+ Oversees the building and customization of moderately complex dashboards, telemetry systems, and alerting mechanisms to proactively monitor components and system health.

** System Design & Architecture - Correctness / Availability:*
* + Oversees the design and implementation of functional and correctness requirements for feature sets and/or systems in new or existing systems.

+ Guides teams to design complex test scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.

+ Directs implementation strategies for data replication and synchronization techniques to maintain data integrity and availability.

** Operational Troubleshooting & Incident Management:*
* + Guides teams to be proactive when diagnosing, debugging, and resolving issues in active components and systems to support ongoing operation.

+ Ensures teams leverage expertise to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.

+ Oversees operational readiness protocol and ensures teams remain knowledgeable of owned components and systems to support effective troubleshooting and performance.

+ Oversees and approves schedules for operational support rotations.

** Compliance & Security:*
* + Oversees implementation of robust security measures to protect data and…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary