Principal Systems Software Engineer
Listed on 2026-08-26
-
Software Development
Embedded Systems/ Firmware/ IoT, DevOps, Unix/Linux, Software Engineer
Job Description
Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems software and platform-management capabilities that power next-generation GPU infrastructure.
This is a hands‑on systems engineering role at the intersection of software, firmware, hardware, and large‑scale cloud infrastructure. You will work on complex GPU and server platforms, owning critical capabilities spanning BMC/service‑processor software, platform management, firmware lifecycle, reliability and serviceability, telemetry, power management, hardware bring‑up, and fleet operations.
You will work closely with silicon, firmware, hardware, compute, and fleet engineering teams, as well as technology and manufacturing partners, to bring new platforms from initial hardware enablement through production deployment and ongoing operation at cloud scale.
This role is well suited for an engineer with deep experience in computer systems, embedded or platform firmware, and hardware/software integration who enjoys solving difficult problems that cross traditional engineering boundaries.
ResponsibilitiesKey Responsibilities:
Design, develop, and maintain complex BMC, service‑processor, and platform‑management software for GPU and server infrastructure.
Develop capabilities for host management, baseboard management, telemetry, platform monitoring, power control, and system diagnostics.
Design and implement secure, reliable firmware‑management and update workflows across complex, multi‑vendor platforms.
Define robust interfaces and integration contracts between platform firmware, hardware components, operating systems, drivers, and higher‑level infrastructure services.
Design, develop, and maintain complex BMC, service‑processor, and platform‑management software for GPU and server infrastructure.
Develop capabilities for host management, baseboard management, telemetry, platform monitoring, power control, and system diagnostics.
Design and implement secure, reliable firmware‑management and update workflows across complex, multi‑vendor platforms.
Define robust interfaces and integration contracts between platform firmware, hardware components, operating systems, drivers, and higher‑level infrastructure services.
Design and implement low-level systems software and firmware using technologies such as C, C , Python, and Bash.
Develop maintainable software for managing, monitoring, diagnosing, and provisioning server and GPU systems.
Build automation and tooling that improves platform provisioning, onboarding, validation, diagnostics, and fleet operations.
Develop and debug advanced platform capabilities involving areas such as RAS, telemetry, power management, high-speed I/O, and chipset/SoC services.
Conduct design and code reviews and help establish strong engineering practices for maintainability, testing, observability, security, and reliability.
Play a leading role in initial board, device, and platform bring‑up for new GPU and server systems.
Diagnose difficult failures spanning hardware, firmware, bootloaders, operating systems, drivers, and platform services.
Use hardware and software diagnostic techniques—including logs, schematics, JTAG, logic analyzers, emulators, and platform instrumentation—to isolate root causes.
Partner with silicon, board, firmware, and manufacturing teams to validate end‑to‑end platform behavior and resolve integration issues.
Turn complex or recurring failures into durable engineering fixes, improved diagnostics, automation, and preventive controls.
Develop reliability, availability, and serviceability (RAS) capabilities for large‑scale GPU and server environments.
Improve telemetry, fault detection, logging, observability, and diagnostics used to identify and resolve platform issues.
Develop and improve power‑control and power‑capping capabilities and related platform instrumentation.
Design systems with fleet‑scale reliability, fault tolerance, secure firmware lifecycle, and operational serviceability in…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).