×
Register Here to Apply for Jobs or Post Jobs. X

Platform Architect

Job in Sacramento, Sacramento County, California, 95828, USA
Listing for: Ddn
Full Time position
Listed on 2026-07-21
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), Systems Engineer, IT Infrastructure, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 140000 - 210000 USD Yearly USD 140000.00 210000.00 YEAR
Job Description & How to Apply Below
Position: Platform Support Architect

DDN is expanding our Enterprise and Sovereign AI Solution offerings, for example Hyperpod - a turnkey NVIDIA AI Data Platform built on DDN Infinia storage, NVIDIA AI Enterprise (NVAIE), and Supermicro reference hardware, optimized for inference and RAG workloads. Our support organization is deep on storage (Infinia, EXAScaler); we are now hiring an AI platform specialist to lead supportability and enablement for the AI side of the stack – NVIDIA AI Enterprise services (NIMs, NeMo, Triton, GPU Operator, licensing), vector databases (initially Milvus), RAG/agentic workflows, and the high‑performance storage and networking fabric that underpins them.

You will be a trusted technical advisor within Support and across OEM and NVIDIA partner teams, combining the mindset of a solutions architect (architecture, reference patterns, PoCs, reusable assets) with that of a L3 support engineer. You’ll help DDN and our partners operate AI Data solutions as a cohesive AI platform, not just a collection of components.

Key Responsibilities
Platform support
  • Act as the primary NVIDIA AI Enterprise and vector database solutions expert for HyperPOD customer environments, bringing deep knowledge of NVAIE services (e.g., NIMs, NeMo, Triton, TensorRT/TensorRT‑LLM, GPU Operator, licensing/NLS) and vector databases (e.g., Milvus) to guide diagnosis, optimization, and solution design.

  • Own complex end‑to‑end triage across GPU, NVAIE services, vector DB, Kubernetes, Docker, high‑speed networking, and Infinia storage, distinguishing product defects from environmental and integration issues.

  • Diagnose and resolve performance bottlenecks in RAG and agentic AI workflows, from model selection and prompt/RAG configuration throughto vector search, GPU utilization, and data access patterns.

  • Collect and interpret logs and telemetry across Linux, containers, Kubernetes, GPU stack, vector DB, and storage/networking; build minimal repros and high‑quality defect reports for escalation to NVIDIA, vector‑DB vendors, OEMs, and internal engineering.

Runbooks, diagnostics, and supportability
  • Author and maintain support triage runbooks and checklists for HyperPOD covering NVAIE services, Milvus/vector DB, GPU stack, Docker, Kubernetes resources, and their interaction with Infinia and the network fabric.

  • Define and validate unified diagnostics bundles that capture the right logs/configs/metrics from all relevant layers (Infinia, GPUs, NVAIE, Milvus, Kubernetes, network) to enable fast problem isolation and high‑signal escalations.

  • Collaborate with observability and tools teams to shape Prometheus/Grafana/ELK/NetQ or equivalent dashboards that surface both platform health and RAG/service‑level metrics (e.g., TTFT, retrieval latency, error rates, throughput).

Enablement, PoCs, and reusable assets
  • Build hands‑on labs and PoCs that mirror customer RAG and agentic AI use cases on HyperPOD, validating supportability and capturing “known good” configurations and troubleshooting patterns.

  • Develop reusable technical assets – implementation guides, best‑practice playbooks, tuning checklists, example architectures – to accelerate time‑to‑value for customers, PS, and Support.

Design feedback, readiness, and cross‑functional leadership
  • Provide structured feedback from early field cases and PoCs into Product Management and Engineering on stack compatibility, upgrade order, rollback constraints, and observability needs for NVAIE, Milvus/cuVS, Infinia, and networking.

  • Collaborate closely with NVIDIA solutions architects, OEM architects, PS, and Support Innovation to align reference architectures and best practices with real‑world support experience.

Required Experience & Skills
Technical
  • 5+ years in Linux‑based infrastructure roles (SRE, MLOps, platform engineering, or L2/L3 support) supporting production systems; 8+ years total technical experience preferred.

  • Strong hands‑on experience with containers and Kubernetes (Docker/containerd, Helm, Operators; debugging pods, Daemon Sets, CSI, CNI, and ingress/load balancers).

  • Demonstrated experience operating GPU‑accelerated workloads in production:

    • NVIDIA GPUs, drivers, CUDA concepts, GPU utilization/perf triage

    • NVIDIA GPU Operator…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary