×
Register Here to Apply for Jobs or Post Jobs. X

Principal Software Engineer, Core Infrastructure

Job in Dover, Kent County, Delaware, 19904, USA
Listing for: Oracle
Full Time position
Listed on 2026-09-26
Job specializations:
  • Software Development
    Software Engineer, DevOps, Cloud Engineer - Software
Job Description & How to Apply Below
** Job Description*
* + Compute & GPU Provisioning:

Experience with BIOS/UEFI, BMC/SMM, Redfish, PLDM, TPM, virtual media, firmware installation and upgrades, and hardware lifecycle management.

+ CPU/GPU & Hardware Systems:
Knowledge of CPU/GPU architectures, PCIe, DMA, NICs, NVMe, device enumeration, hardware bring-up, burn-in validation, and firmware debugging.

+ Bare-Metal & OS Provisioning:

Experience with PXE/iPXE, DHCP/DNS, Linux networking, OS imaging, and secure fleet-scale provisioning and automation.

+ Leads development and begins architecting scalable, container-based services that build, validate, and monitor liquid-cooled GPU infrastructure across factory environments.

+ Develops secure infrastructure, control-plane workflows, and data-plane test capabilities that orchestrate manufacturing test life cycles, operator workflows, quality gates, fleet status, and business reporting for platforms including GB200, GB300, VR, MI355, and MI455.

+  Automates and maintains GPU test validation; manages consistent firmware, software, and hardware configuration; and develops repair and triage capabilities that identify root causes and guide recovery.

+ Partners with Supply Chain Operations, Hardware Development, external manufacturing partners, data-center operations, NVIDIA, and AMD to resolve issues before racks ship.

+ Establishes manufacturing yield, throughput, quality, and deployment-readiness metrics that reduce rework and downstream failures, improve data-center ingestion, and accelerate reliable hyperscale AI infrastructure delivery.

** Responsibilities*
* ** Responsibilities*
* + Lead the design, implementation, and ongoing evolution of core distributed systems and data-plane services at hyperscale.

+ Define scalability, elasticity, durability, and availability requirements for owned components and ensure designs meet them.

+ Optimize high-throughput data paths for large-scale retrieval, storage, and processing using distributed state, replication, and synchronization patterns.

+ Design fault-tolerant systems that support in-service updates through redundancy, automatic failover, and recovery-oriented design.

+ Apply sound distributed-systems tradeoffs for network partitions and reliability, including load shedding, throttling, rate limiting, retries, and timeouts.

+ Establish service-level objectives, key performance indicators, telemetry, dashboards, and proactive alerting for critical systems.

+ Design and lead performance, load, fault-injection, and brownout testing to validate correctness, resilience, and operational readiness.

+ Lead production incident diagnosis and recovery, guide root-cause analysis, and mentor engineers in operational excellence.

+ Build and improve Infrastructure as Code and operational automation that enable safe patching, updates, rollbacks, and change management.

+ Apply robust security controls and remediation practices for multi-tenant cloud infrastructure, including encryption, access controls, and compliance readiness.
** Qualifications*
* + Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.

+ 7+ years of professional software-engineering experience, with demonstrated impact on large-scale distributed systems or cloud infrastructure.

+ Strong experience designing and operating highly available, scalable, fault-tolerant distributed systems.

+ Proficiency in one or more object-oriented or systems programming languages, such as Java, C++, C#, or Go.

+ Deep understanding of distributed-systems design, data structures, algorithms, operating systems, networking, and secure software-development practices.

+

Experience with system-level test automation, performance/load testing, reliability…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary