×
Register Here to Apply for Jobs or Post Jobs. X

Senior Software Engineer, Core Infrastructure

Job in Seattle, King County, Washington, 98127, USA
Listing for: Oracle Corporation
Full Time position
Listed on 2026-09-06
Job specializations:
  • IT/Tech
Job Description & How to Apply Below
Job Description

Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers' implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery-oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry;

authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.

Oracle Cloud Infrastructure Workflow is a Tier 0 service that is critical to the smooth functioning of ALL basic OCI services by enabling their execution of distributed, multi-step work in a fault tolerant manner. An engineer on this team is responsible for the development of features making the platform more resilient and efficient including the launch of a brand new V2 version, as well as the operational excellence of the service, as we scale to meet OCI's exponential growth needs.

Responsibilities

Key Responsibilities
System Design & Architecture - System Scalability:
-Implements and contributes to the development for components of distributed systems that support horizontal and vertical scaling including leveraging distributed state management tools.
-Optimizes code and/or systems for large-scale data processing in large-scale systems.
-Implements scalability requirements for assigned components and reviews implementation of team members.
-Leverages components of data plane platforms to handle large-scale data retrieval, storage, and processing.
-Implements performance and load testing.
System Design & Architecture - System Reliability Design:
-Collaborates with team to build fault-tolerant components capable of withstanding in-service updates by implementing redundancy, replication, and automatic failover mechanisms.
-Applies recovery oriented computing principles to design components that effectively handle service disruptions.
-Implements retry mechanisms, circuit breakers, and timeouts to help handle network unreliability.
System Design & Architecture - System Reliability Performance:
-Implements tests and alarm configurations to proactively detect and address issues/failures.
-Supports efforts to recover from failures by drafting and executing runbooks and operational procedures.
-Builds and customizes dashboards, telemetry systems, and alerting mechanisms to monitor component health.
System Design & Architecture - Correctness / Availability:
-Designs and implements functional requirements and testing for assigned features within an existing system.
-Implements tests scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.
-Implements standard data replication and synchronization techniques to maintain data integrity and availability.
Operational Troubleshooting & Incident Management:
-Diagnoses, debugs, and resolves issues in system components to support ongoing operation.
-Implements basic strategies to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
-Designs and implements automation scripts and tooling used to troubleshoot operational issues.
-Participates in operational support rotations, assisting in incident responses and root cause investigations.
Compliance & Security:
-Applies advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
-Implements remediation plans to continuously improve security.
-Collaborates with the team to ensure cloud infrastructure complies with relevant industry standards and regulations and that documentation is up-to-date
Automation & Change Management:
-Maintains automation scripts and tools (e.g., Infrastructure as Code (IaC)) for managing cloud infrastructure.
-Adheres to change management plans for patching, updating, and rolling back applications.

Core Responsibilities
Planning & Execution:
-Track timelines with minimal supervision, ensuring work is completed in a timely manner and is in alignment with project requirements.
-Prioritize and adjust work as resources or timelines change, with some guidance
Collaboration & Partnership:
-Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
Problem Solving:
-Independently identifies and addresses…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary