More jobs:
AI Systems Engineer; OCI/AI Infrastructure
Job in
Jackson, Hinds County, Mississippi, 39203, USA
Listed on 2026-08-29
Listing for:
Oracle
Full Time
position Listed on 2026-08-29
Job specializations:
-
IT/Tech
Systems Engineer, AI Engineer (Applied/Software)
Job Description & How to Apply Below
** Job Description*
* Oracle Hardware Platform Development Engineering is seeking a highly driven
** AI Systems Engineer
** to evaluate and characterize next-generation GPU and AI accelerator platforms for Oracle Cloud Infrastructure (OCI). This is a hands-on engineering role focused on bringing up new hardware platforms, enabling AI training and inference software stacks, running representative workloads, and analyzing system performance under real operating conditions.
The engineer will identify whether workloads are
** HBM/memory-bandwidth, compute, scale-up, or scale-out bound** , while characterizing power, thermals, memory behavior, utilization, scaling, and performance efficiency. Working directly in the lab, you will debug hardware/software integration issues, design and execute experiments, and develop data-driven insights that explain system behavior beyond benchmark results.
A key part of the role is
** comparative architecture analysis
** across GPUs and emerging AI accelerators. You will evaluate architectural tradeoffs and translate performance findings into clear, actionable recommendations on which platforms are best suited for specific AI training and inference workloads. You will work closely with internal hardware and software teams as well as technology partners to help shape Oracle's next generation of high-performance AI infrastructure.
Position Overview:
This position is ideal for someone who loves deep systems engineering, debugging complex hardware-software interactions, and optimizing performance at every layer of the ML stack. You will play a pivotal role in enabling the training and deployment of next-generation LLMs and generative AI models.
** Responsibilities*
* Required Qualifications
+ Solid knowledge of AI / GPU or/and AI/CPU platform architecture and their capabilities.
+
Experience with the architecture, design, and implementation of modern server platforms consisting of multiple architectures and vendors, including x86 and ARM server architectures.
+ Strong communications skills and ability to clearly communicate complex technical issue across engineering disciplines as well as clearly and succinctly articulate issues for executives.
+ Experience and understanding of the latest high-speed busses and interconnect used in modern Compute and AI platforms. Familiarity with their startup connectivity and operational robustness as well as performance metrics.
+ Debugging & Reliability:
Troubleshoot complex hardware-software interaction issues, including vLLM compilation failures on ROCm, CUDA memory leaks, distributed runtime failures, and kernel-level inconsistencies.
+ Profiling & Performance Analysis:
Conduct detailed profiling of compilation graphs, training workloads, and runtime execution to optimize performance and eliminate bottlenecks.
Preferred Qualifications
+ Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.
+ Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).
+ Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level.
+ Hands-on experience maintaining or building ML training stacks involving CUDA, ROCm, NCCL, XLA, or similar technologies.
+ Experience in benchmarking AI workloads across different architectures.
+ Background in working with the large scale clusters
+ Good understanding on DL frameworks internal PyTorch, Tensor Flow, JAX, and Ray
Disclaimer:
** Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.*
* ** Range and benefit information provided in this posting are specific to the stated locations only*
* US:
Hiring Range in USD from: $96,800 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×