×
Register Here to Apply for Jobs or Post Jobs. X

Storage Platform Engineer; AI Storage; Remote

Remote / Online - Candidates ideally in
Boise, Ada County, Idaho, 83701, USA
Listing for: Greenhouse Software, Inc.
Remote/Work from Home position
Listed on 2026-10-06
Job specializations:
  • Software Development
    Data Engineering
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below
Position: Staff Storage Platform Engineer (AI Storage) - Remote

Remote

About Radian Arc.

We're specialists in outcome-optimized AI infrastructure - deploying, orchestrating and monetizing GPU compute where data, users and demand actually meet: inside telco networks, at the edge, and in core data centers.

Not a generic AI platform. Not a consultancy. We’re the bridge between raw silicon and real-world results.

What impact you will have

Mission:
Design, build, and operate the AI storage layer powering large-scale GPU infrastructure, enabling datasets, model artifacts, checkpoints, and inference state to be delivered to compute clusters with extremely high throughput and predictable latency.

You will play a key role in architecting and evolving the storage platform across edge and core deployments, supporting the full lifecycle of AI workloads including distributed inference, fine-tuning, and large-scale model training. The role spans multiple storage architectures used across the platform, including hyperconverged storage currently based on Stor Pool, local NVMe storage for latency-sensitive workloads and edge deployments, and disaggregated AI storage platforms such as VAST Data and Weka.

As the first dedicated storage platform role in the organization, this position combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands-on execution across storage design, deployment, performance engineering, troubleshooting, platform integration, and operational improvement.

A key responsibility of this role is designing and optimizing the storage architecture underlying distributed inference stacks such as NVIDIA Dynamo, llm-d, or similar inference orchestration frameworks. This includes ensuring that storage systems efficiently support inference workloads through optimized dataset access, model artifact distribution, checkpoint handling, and KV-cache persistence. You will design scalable storage systems capable of feeding thousands of GPUs while balancing throughput, latency, resilience, and cost efficiency, and work closely with compute, networking, and platform engineering teams to ensure seamless integration with the platform orchestration layer.

Because this is currently the primary storage platform role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of long-term design, standards, cross-team influence, and platform direction, while also directly executing critical storage work that, in a larger organization, would be distributed across multiple engineers.
This is a fully remote role, we will consider relevant candidates in all locations.

What you’ll need

Core Experience
  • Strong hands-on experience designing and operating distributed storage systems for high-performance compute environments.
  • Proven experience designing storage architectures for large-scale AI inference or training platforms, including dataset distribution, checkpointing, and KV-cache storage patterns.
  • Deep knowledge of the Linux storage and I/O stack.
  • Strong understanding of AI workload data access patterns.
  • Experience optimizing storage for GPU-accelerated workloads.
  • WEKA Data Platform (an enterprise high-performance storage system for AI and HPC)
  • Familiarity with Kubernetes storage integrations such as CSI.
  • Experience owning both architecture and direct implementation in lean or fast-scaling environments is strongly preferred.
Advanced AI Storage Expertise

The candidate should have deep expertise in designing and operating storage platforms optimized for GPU-heavy environments and distributed AI workloads.

This includes a strong understanding of how training, fine‑tuning, and inference systems interact with storage, and how storage architecture affects throughput, latency, concurrency, checkpoint…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary