×
Register Here to Apply for Jobs or Post Jobs. X

Senior SoC Architect, RAS

Job in Northern, Floyd County, Kentucky, USA
Listing for: NVIDIA Corporation
Full Time position
Listed on 2026-09-04
Job specializations:
  • Engineering
    Systems Engineer, Hardware Engineer, Test Engineer
Salary/Wage Range or Industry Benchmark: 184000 - 356500 USD Yearly USD 184000.00 356500.00 YEAR
Job Description & How to Apply Below
Location: Northern

We are now looking for a Senior Hardware Architect for our Tegra System-on-Chips (SoC) focused on Reliability, Availability, and Serviceability (RAS). Do you want to be part of the Artificial Intelligence (AI) revolution and help define resilient computing platforms for datacenters, autonomous vehicles, edge systems, and other high-reliability applications? We are looking for an exceptional SoC architect to help define, drive, and deliver RAS hardware architecture across advanced CPUs and SoCs, from early architectural concepts through design implementation, verification, validation, and production readiness.

This position offers the opportunity to have real impact in a dynamic, technology-focused company developing state-of-the-art processor and system architectures at the forefront of machine learning, autonomous vehicles, high-performance computing, and edge computing. You will work with world-class systems architects, RAS experts, design teams, verification teams, validation teams, firmware teams, and software partners to define end-to-end hardware RAS features that improve system resiliency, observability, debuggability, error containment, recovery, and serviceability.

Space and radiation-aware design are important areas of interest for this role, including understanding how radiation effects can influence SoC reliability, but the primary focus is broad SoC RAS architecture and driving features successfully through the product development flow.

What you'll be doing:

Define and drive SoC-level RAS hardware architecture across CPUs, interconnects, memory systems, IOs, safety islands, firmware interfaces, and platform-level components.

Own RAS features from concept through architecture specification, micro-architecture alignment, RTL implementation support, design verification, silicon validation, debug, and production readiness.

Develop architectural requirements for fault detection, correction, containment, isolation, telemetry, error reporting, recovery, graceful degradation, serviceability, and diagnostic observability.

Work closely with design, verification, validation, firmware, software, and platform teams to ensure RAS features are implementable, verifiable, debuggable, and aligned with system-level requirements.

Understand the broader SoC architecture and identify how RAS mechanisms interact with performance, power, reset flows, clocks, memory-hierarchy, interconnect behavior, firmware-visible controls, and platform software.

Create hardware specifications, architectural requirements, error-handling flows, design guidance, test plans, and architectural models in SystemC, C/C++, Python, or other relevant modeling environments where applicable.

Plan and review verification and validation strategies for RAS mechanisms, including error injection, recovery validation, coverage analysis, resiliency modeling, and cross-functional architecture reviews.

Assist in failure analysis and silicon debug for lab, post-silicon, production, and field findings; develop diagnostic screens and localization methods for latent, intermittent, and environment-sensitive failures.

Apply RAS architecture principles to high-reliability deployment environments, including space-aware and radiation-aware use cases where single-event effects, memory corruption, logic corruption, or cumulative radiation exposure may impact system reliability.

Follow industry standards and best practices related to RAS, functional safety, semiconductor reliability, debuggability, verification, validation, and silicon testing.

Patent novel hardware architecture techniques that improve system resiliency, observability, serviceability, and recovery.

What we need to see:

MS or PhD degree in computer engineering, electrical engineering, or equivalent experience.

At least 8+ years of SoC architecture, design, verification, reliability, silicon validation, or related hardware development experience.

Strong understanding of Reliability, Availability, and Serviceability (RAS) in the SoC context, including fault detection, correction, containment, telemetry, recovery, degradation modes, debug visibility, and serviceability mechanisms.

Experience defining and driving hardware architecture features through the full development lifecycle, including architecture definition, design implementation, verification planning, validation, debug, and production readiness.

Strong understanding of overall SoC architecture and the ability to reason across micro-architecture, full-chip integration, firmware interfaces, software-visible behavior, platform flows, and customer use cases.

Meaningful industry expertise in one or more SoC architecture areas such as RAS, safety, debug, clocks, resets, interconnects, memory controllers, IO technologies, platform integration, firmware-visible error handling, or diagnostic infrastructure.

Hands-on experience with design verification, silicon validation, fault injection, coverage analysis, resiliency modeling, diagnostic development, or reliability…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary