Senior Core Infrastructure Engineer (OCI Object Storage
Job in
Seattle, King County, Washington, 98127, USA
Listed on 2026-08-15
Listing for:
Oracle Corporation
Full Time
position Listed on 2026-08-15
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Job Description & How to Apply Below
Job Description
Are you interested in building large-scale distributed infrastructure for the cloud? Oracle's Cloud Infrastructure team is building Infrastructure-as-a-Service technologies that operate at high scale in a broadly distributed multi-tenant cloud environment. Our customers run their businesses on our cloud, and our mission is to provide them with industry leading compute, storage, networking, database, security, and an ever expanding set of foundational cloud-based services.
As part of this effort, the Object Storage Service team is looking for hands-on engineers with expertise and passion in solving difficult problems in distributed systems, large scale storage, and highly available services. If this is you, you can be part of the team that drives the best-in-class Object Storage Service into the next phase of its development. These are exciting times for the service - we are growing fast, and delivering on innovative, enterprise class features to satisfy the most demanding big data and enterprise workloads for our customers.
An engineer at any level can have significant technical and business impact.
As a senior engineer, you will own the software design and development for major components and features of the Object Storage Service. You should be both a rock-solid coder and with familiarity of distributed systems. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn.
Responsibilities
Key Responsibilities
System Design & Architecture - System Scalability:
-Implements and contributes to the development for components of distributed systems that support horizontal and vertical scaling including leveraging distributed state management tools.
-Optimizes code and/or systems for large-scale data processing in large-scale systems.
-Implements scalability requirements for assigned components and reviews implementation of team members.
-Leverages components of data plane platforms to handle large-scale data retrieval, storage, and processing.
-Implements performance and load testing.
System Design & Architecture - System Reliability Design:
-Collaborates with team to build fault-tolerant components capable of withstanding in-service updates by implementing redundancy, replication, and automatic failover mechanisms.
-Applies recovery oriented computing principles to design components that effectively handle service disruptions.
-Implements retry mechanisms, circuit breakers, and timeouts to help handle network unreliability.
System Design & Architecture - System Reliability Performance:
-Implements tests and alarm configurations to proactively detect and address issues/failures.
-Supports efforts to recover from failures by drafting and executing runbooks and operational procedures.
-Builds and customizes dashboards, telemetry systems, and alerting mechanisms to monitor component health.
System Design & Architecture - Correctness / Availability:
-Designs and implements functional requirements and testing for assigned features within an existing system.
-Implements tests scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.
-Implements standard data replication and synchronization techniques to maintain data integrity and availability.
Operational Troubleshooting & Incident Management:
-Diagnoses, debugs, and resolves issues in system components to support ongoing operation.
-Implements basic strategies to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
-Designs and implements automation scripts and tooling used to troubleshoot operational issues.
-Participates in operational support rotations, assisting in incident responses and root cause investigations.
Compliance & Security:
-Applies advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
-Implements remediation plans to continuously improve security.
-Collaborates with the team to ensure cloud infrastructure complies with relevant industry standards and regulations and that documentation is up-to-date
Automation & Change Management:
-Ma…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×