Principal Core Infrastructure Engineer; OCI Object Storage
Listed on 2026-07-18
-
Software Development
Backend Developer, Cloud Engineer - Software, Software Engineer, Software Architect
Are you interested in building large-scale distributed infrastructure for the cloud? Oracle’s Cloud Infrastructure team is building Infrastructure-as-a-Service technologies that operate at high scale in a broadly distributed multi-tenant cloud environment. Our customers run their businesses on our cloud, and our mission is to provide them with industry leading compute, storage, networking, database, security, and an ever expanding set of foundational cloud-based services.
As part of this effort, the Object Storage Service team is looking for hands‑on engineers with expertise and passion in solving difficult problems in distributed systems, large scale storage, and highly available services. If this is you, you can be part of the team that drives the best‑in‑class Object Storage Service into the next phase of its development. These are exciting times for the service - we are growing fast, and delivering on innovative, enterprise class features to satisfy the most demanding big data and enterprise workloads for our customers.
An engineer at any level can have significant technical and business impact.
As a senior engineer, you will own the software design and development for major components and features of the Object Storage Service. You should be both a rock‑solid coder and with familiarity of distributed systems. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn.
Responsibilities- Leads development and begins architecting components of scalable, elastic distributed systems.
- Defines and enforces scalability requirements for owned components; optimizes code and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms for large‑scale retrieval, storage, and processing.
- Designs fault‑tolerant, in‑service‑upgradable systems using redundancy, replication, failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting to handle network unreliability while meeting SLOs.
- Establishes KPIs and telemetry; builds proactive dashboards and alerts; and designs complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability.
- Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness.
- Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation that enable safe patching, updates, and rollbacks within change‑management plans.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).