Failure Analysis Manager
Listed on 2026-10-05
-
Engineering
Quality Engineering -
Quality Assurance - QA/QC
Quality Engineering
Title: Senior Manager, Failure Analysis Engineering
Location: Georgetown Texas
Duration:Direct Hire
Work Requirements: US Citizen, GC Holders or Authorized to Work in the U.S.
Role Overview
The Senior Manager, Failure Analysis Engineering serves as the technical authority for system-level failure analysis and product reliability across the full manufacturing lifecycle, from New Product Introduction (NPI) through High Volume Manufacturing (HVM).
This role leads the identification of complex failure mechanisms, defines structured root cause methodologies, and drives cross-functional resolution to improve product quality, manufacturing yield, and long-term reliability
.
The position requires deep expertise in server hardware architectures, failure physics, and data-driven analysis
, combined with the ability to influence engineering, quality, and manufacturing organizations.
- NPI, including DVT / PVT readiness
- Production, including L6, L10, and system-level testing
- Field and customer returns, including RMA and DOA
- CPU
- Memory
- Storage
- Power
- Networking
- Thermal systems
- Test Engineering
- Product Engineering
- Quality, including PQE / MQE
- Supplier Engineering
- Manufacturing
- Data-driven reliability and failure trend analysis
- Influence on product design
- Test coverage improvements
- Manufacturing process improvements
- Lead complex failure analysis (FA) and root cause analysis (RCA) to identify system-level failure mechanisms across server platforms.
- Define and standardize failure analysis methodologies, tools, processes, and best practices
. - Drive reliability strategy and influence NPI readiness
, including DFR, DFT, and test coverage. - Establish failure trend analysis across yield, escapes, and field returns to enable data-driven decision-making.
- Serve as an escalation point for critical quality issues and lead cross-functional technical problem solving.
- Drive corrective and preventive actions across design, test, and manufacturing to eliminate repeat failures.
- Improve test effectiveness, reduce NTF (No Trouble Found) loops
, and strengthen feedback loops into engineering. - Mentor engineers and elevate failure analysis capabilities across the organization.
- Identify systemic failure drivers and develop technical strategies to improve product reliability and manufacturing performance.
- Deep knowledge of server hardware architectures
, including:- CPU
- Memory
- Storage
- Power
- Networking
- Strong expertise in failure analysis methodologies and root cause analysis techniques
. - Experience with system-level debugging
, including:- Electrical issues
- Firmware issues
- Hardware/software integration issues
- Strong statistical analysis and data interpretation skills, including:
- Manufacturing yield
- Reliability Failure trends
- Understanding of manufacturing test flows, including:
- L6
- L10
- System-level testing
- Experience with technical and data analysis tools, including:
- Oscilloscopes
- Logic analyzers
- Diagnostic tools
- Python
- SQL
- Power BI or equivalent data visualization/analysis tools
- Experience with GPU systems and liquid cooling
. - Experience in hyperscale manufacturing environments.
- Automation experience.
- Six Sigma certification
. - Experience supporting hyperscale or data center server environments
. - Knowledge of reliability modeling, including:
- Weibull analysis
- MTBF
- HALT
- HASS
- Exposure to DFX methodologies
, including:- DFR
- Design for Reliability - DFT
- Design for Test - DFM
- Design for Manufacturing
- DFR
- Experience automating failure analysis workflows and data pipelines
. - Experience working with suppliers and supporting component-level failure analysis
.
- Bachelor's or Master's degree in:
- Electrical Engineering
- Computer Engineering
- or a related technical field
- 10+ years of experience in:
- Failure analysis
- Reliability engineering
- System debugging
- Hardware engineering
- Related technical disciplines
- Proven experience supporting NPI through HVM transitions in complex hardware systems.
- Demonstrated track record of solving complex, cross-domain technical problems.
- Strong ability to influence engineering, quality, manufacturing, and supplier organizations.
- First 6 Months
- Establish a structured failure analysis framework and RCA methodology across programs.
- Identify the top systemic failure drivers and implement corrective actions.
- Improve failure containment and reduce repeat issues and NTF rates
. - Build strong cross-functional alignment across Test Engineering, Quality, Product…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).