×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Failure Analysis Engineer

Job in Newark, Alameda County, California, 94560, USA
Listing for: sghcorp.com
Full Time position
Listed on 2026-07-30
Job specializations:
  • Engineering
    Systems Engineer
Salary/Wage Range or Industry Benchmark: 145000 - 165000 USD Yearly USD 145000.00 165000.00 YEAR
Job Description & How to Apply Below

Select how often (in days) to receive an alert:

At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide.

Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.

Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.

At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.

Overview

We are seeking a seasoned and highly skilled Senior DRAM Failure Analysis Engineer with extensive expertise in high-speed, high-capacity DRAM modules. The ideal candidate will have 5+ years of experience in DRAM module reliability testing, high-speed signal integrity analysis, and failure analysis to ensure optimal performance in mission-critical environments. This leadership role involves owning and advancing our burn-in methodologies, spearheading high-speed testing strategies, and driving system-level failure analysis to guarantee the reliability of DDR4/DDR5 based memory solutions.

You will work closely with cross-functional teams to define and enhance product performance, quality, and reliability, while also mentoring junior engineers.

Responsibilities
  • Lead and conduct complex failure analysis (FA) on DRAM modules and memory subsystems, utilizing high-speed signal integrity tools and oscilloscopes.
  • Drive the determination of failure root causes at the chip, module, and system level, presenting findings to technical and leadership teams.
  • Architect and execute debug strategies for system-level failures by analyzing memory controller interactions, BIOS tuning, and DIMM register settings in server and cloud environments.
  • Investigate and resolve critical performance bottlenecks, intermittent failures, and memory errors caused by power integrity (PI), signal integrity (SI), and thermal stress.
  • Lead the development and implementation of advanced stress test strategies for high-performance DRAM modules.
  • Mentor junior engineers in best practices for high-frequency waveform analysis, signal integrity debugging, and root cause analysis.
  • Generate and present detailed reports (8D) summarizing failure analysis findings, root causes, and strategic recommendations for product and process improvements.
  • Act as a technical lead in cross-functional teams to enhance product design, quality, and manufacturability.
Qualifications
  • Bachelor’s degree in Electrical Engineering or a related field;
    Master’s degree is a plus.
  • 5+ years of experience in DRAM module burn-in, stress testing, and failure analysis.
  • Expert-level understanding of high-speed memory interfaces (DDR4, DDR5, HBM) and advanced SI/PI concepts.
  • Proven track record of complex problem-solving and root cause failure analysis in a high-performance computing environment.
  • Demonstrated mastery of server memory modules and system-level debugging in Linux and Windows environments.
  • Expertise with high-frequency test equipment such as oscilloscopes.
  • AI & Automation Fluency:
    Strong proficiency in applying modern generative AI tools and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary