×
Register Here to Apply for Jobs or Post Jobs. X

Principal AI SoC Runtime Software Architect

Job in Santa Clara, Santa Clara County, California, 95050, USA
Listing for: Auradine
Full Time position
Listed on 2026-08-15
Job specializations:
  • Software Development
    Software Architect, Software Engineer, C++ Developer
Job Description & How to Apply Below

Principal Ai Soc Runtime Software Architect

We are looking for a Principal AI SoC Runtime Software Architect to own the software architecture that turns Velaura's heterogeneous AI SoC into a coherent, high-performance execution platform.

This role will define and lead development of the end-to-end runtime spanning sensor ingest, preprocessing, AI inference, postprocessing, and delivery of results to robotics and other physical and edge AI applications. The runtime will coordinate execution and data movement across the SoC's CPU cores, AI accelerator, vision and multimedia engines, and other embedded processors. The complete runtime must coordinate heterogeneous workloads, manage ownership, synchronization, and safe reuse of shared data buffers, minimize data movement, provide predictable low-latency execution, recover from failures, and expose cohesive APIs and observability to applications and SDK components.

The ideal candidate combines deep runtime and systems-software expertise with a strong understanding of heterogeneous compute, Linux kernel and driver interfaces, DMA and shared-memory architectures, and performance-sensitive AI or multimedia pipelines. This will be a hands-on principal architect and technical lead who directs engineers across runtime, kernel, driver, firmware, multimedia, and SDK development.

Responsibilities
  • Define the SoC-wide execution model for coordinating workloads across heterogeneous compute and media engines, including dependency management, resource arbitration, priority and QoS, and concurrent pipeline behavior.

  • Set the architecture and technical direction for the AI inference runtime, including compiled-model execution, integration with industry-standard AI execution frameworks, and interfaces to applications and the broader Velaura SDK.

  • Own the end-to-end dataflow architecture for sensor-to-application pipelines, ensuring that camera, media, preprocessing, inference, and postprocessing components operate as an efficient and coherent system.

  • Define the SoC-wide memory and buffer-sharing architecture across user space, the kernel, and heterogeneous hardware engines, establishing clear ownership, coherency, isolation, synchronization, and lifecycle semantics.

  • Define the division of responsibility and interface contracts among the runtime, kernel drivers, firmware, and hardware engines, including execution, completion, telemetry, fault management, and recovery semantics.

  • Partner with the compiler team to define the compiler–runtime contract, ensuring compiled artifacts contain the metadata and execution information needed for the runtime to load, validate, execute, profile, and maintain compatibility across releases.

  • Establish a system-wide observability and performance architecture that correlates behavior across software and hardware layers and enables optimization against latency, throughput, bandwidth, power, utilization, and predictability goals.

  • Define the runtime resilience and validation architecture, including fault-containment and recovery policies, architecture-level acceptance criteria, and qualification across correctness, concurrency, compatibility, performance, and sustained workloads.

Required Qualifications
  • Extensive experience designing and building production runtime systems, embedded middleware, multimedia frameworks, or other performance-critical systems software.

  • Strong C/C++ programming skills and demonstrated ability to architect and contribute hands-on to production runtime software spanning application-facing APIs, user-space libraries, and low-level driver, firmware, and hardware interfaces.

  • Strong understanding of heterogeneous and asynchronous execution, including command submission, queues, events, dependencies, synchronization, concurrency, scheduling, and resource management.

  • Strong understanding of device memory, DMA, IOMMU/SMMU, cache coherency, memory mapping, shared buffers, buffer lifetimes, and kernel/user-space memory interfaces.

  • Experience optimizing end-to-end data movement and execution across multiple hardware engines rather than focusing solely on individual kernels or accelerator performance.

  • Experience designing stable runtime APIs with well-defined compatibility, versioning, error handling, diagnostics, and recovery behavior.

  • Demonstrated ability to debug complex cross-layer correctness and performance problems using disciplined, data-driven methods, profiling, and tracing.

  • Demonstrated technical leadership across component and organizational boundaries, including translating system requirements into clear architectures, interfaces, implementation guidance, and validation strategies.

Preferred Qualifications
  • Direct experience developing or extending AI inference runtimes or execution providers, such as ONNX Runtime, TensorRT-like runtimes, OpenVINO, Tensor Flow Lite delegates, Qualcomm QNN/SNPE, TVM runtimes, or comparable systems for NPUs, GPUs, DSPs, or other accelerators.

  • Linux kernel development or upstream contribution experience involving…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary