Storage Systems Engineer
Listed on 2026-07-14
-
IT/Tech
Systems Engineer, Data Engineering, IT Infrastructure, Cloud Computing: Infrastructure & Operations
Job Description
This is a two-year contract of employment, inclusive of benefits. The Academic Research Services team at UCSF is seeking a Storage Systems Engineer (SYS ADM
4) to serve as a technical resource in the design, deployment, and operation of large-scale research storage and data infrastructure. This role will work in close partnership with the Senior Research Dev Ops Engineer to support UCSF’s evolving research ecosystem, including CoreHPC, the Research Analysis Environment (RAE), and large institutional storage initiatives.
This position is primarily responsible for architecture, implementation, and lifecycle management for the Facility for Advanced Computing (FAC), storage and systems, including support for large storage environments, NSF-funded infrastructure, and OS Nexus–aligned data platforms. The role ensures seamless integration between storage systems and the CoreHPC compute cluster, enabling performant, reliable, and scalable data access for AI, data science, and computational research workloads.
TheStorage Systems Engineer Will
- Work with the lead to continue supporting the design and evolution of storage architecture across on-prem and hybrid environments, including VAST, parallel file systems, and enterprise storage platforms.
- Develop and maintain data movement strategies and tooling (e.g., rsync, rclone, Globus, SMB workflows) to support large-scale data ingestion, migration, and lifecycle management.
- Ensure tight integration between storage and HPC compute systems, optimizing throughput, latency, and reliability for distributed workloads.
- Support and scale storage systems backing major institutional initiatives (FAC storage, OS Nexus integration).
- Collaborate closely with Dev Ops, networking, and security teams to deliver cohesive research infrastructure solutions.
- Design and implement monitoring, performance tuning, and capacity planning strategies for storage and data systems.
- Troubleshoot complex issues across storage, networking, and compute boundaries.
- Participate in system upgrades, migrations, and expansion efforts with minimal disruption to researchers.
- Provide guidance to researchers on data organization, transfer strategies, and performance optimization.
- Evaluate and recommend emerging storage technologies and architectures.
Academic Research Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository of HIPAA compliant clinical and life sciences data and a centralized, secure, professionally managed infrastructure for the storage and management of research data. ARS empowers medical scientific investigations by offering secure computing environments, data capture, management and analysis tools, and support services which meet researchers’ needs.
The Research Infrastructure team of the Academic Research Service focuses on large scale research platform support, high performance computational and storage services for UCSF researchers so they can address complex computational, AI, and data science problems.
- Design, deploy, and operate large-scale storage systems, including ZFS, VAST and parallel file systems.
- Define standards for performance, redundancy, and scalability on ZFS file systems.
- Lead the evolution of institutional storage platforms, including FAC storage environments. Primarily ZFS.
- Manage Active Directory integration with storage and compute systems.
- Architect and execute large-scale data migrations.
- Develop and maintain data movement workflows using tools such as rsync, rclone, and Globus.
- Optimize data transfer processes across storage and compute environments.
- Integrate storage systems with the CoreHPC compute cluster.
- Optimize I/O performance for AI, machine learning, and HPC workloads.
- Support efficient data access patterns for distributed and scheduled workloads.
- Manage and configure VMWare, Bare Metal Servers.
- Implement monitoring, alerting, and capacity planning for storage systems & operating systems like Linux/Windows.
- Troubleshoot issues across storage, network, and compute infrastructure.
- Perform system maintenance, patching, and lifecycle management.
- Advise researchers on data workflows…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).