Sr. HPC and Data Storage Engineer
Listed on 2026-08-26
-
IT/Tech
Unix/Linux, Cloud Computing: Infrastructure & Operations, IT Infrastructure, Systems Administrator
Sr. HPC and Data Storage Engineer (Finance)
Position SummaryThe J. Craig Venter Institute (JCVI) is a global nonprofit genomic research organization dedicated to advancing the science of genomics, human health, and environmental biology through innovative research and technology. The Sr. High Performance Computing (HPC) & Data Storage Engineer will support, maintain, and optimize computing and data storage infrastructure supporting bioinformatics, genomics, research computing, and scientific applications.
This position is responsible for HPC infrastructure, large-scale data storage, Ceph, Linux systems administration, cloud infrastructure, virtualization, infrastructure automation, backup, and disaster recovery. The HPC and Data Storage Engineer will work closely with bioinformatics scientists, software engineers, laboratory personnel, and technical leadership to maintain reliable, scalable, and high-performance computing environments capable of supporting large and rapidly growing scientific datasets.
This position will require deep Linux and storage expertise and experience supporting HPC or other data-intensive computing environments. AWS cloud and Kubernetes experience are highly desirable as the organization's infrastructure continues to evolve.
Position ResponsibilitiesHPC & Research Computing
- Support, maintain, and optimize high-performance computing (HPC) infrastructure supporting bioinformatics, genomics, and data-intensive scientific workloads.
- Administer HPC compute nodes and supporting Linux infrastructure.
- Monitor and optimize system performance, resource utilization, and availability.
- Support scientific software and computational environments used for genomic sequencing and analysis.
- Troubleshoot performance issues involving compute, storage, and applications.
- Support workload scheduling environments such as Slurm or similar HPC schedulers.
- Partner with bioinformatics and software engineering teams to optimize computational workflows and infrastructure.
- Participate in HPC capacity planning and infrastructure architecture decisions.
- Administer, maintain, and optimize large-scale storage infrastructure supporting HPC and scientific computing workloads.
- Deploy, administer, monitor, and troubleshoot Ceph distributed storage environments.
- Manage high-capacity research storage, Linux file systems, and shared storage resources.
- Support block, object, file, parallel, and distributed storage technologies.
- Monitor and optimize storage capacity, throughput, latency, I/O performance, availability, and system health, and troubleshoot complex storage performance issues.
- Perform capacity planning to support rapidly growing scientific datasets.
- Support storage migrations, upgrades, lifecycle management, and high-speed data movement.
- Maintain backup, replication, archival, and disaster recovery solutions, including periodic recovery testing.
- Administer enterprise Linux systems including Red Hat Enterprise Linux, Rocky Linux, Ubuntu, Amazon Linux, or similar distributions.
- Install, configure, patch, upgrade, monitor, and maintain physical and virtual Linux servers.
- Configure file systems, storage mounts, system resources, and services.
- Troubleshoot complex operating system, hardware, storage, and application issues.
- Perform Linux performance analysis and tuning for compute- and data-intensive workloads.
- Administer and support virtualization technologies.
- Support AWS and hybrid cloud infrastructure, including cloud-hosted Linux compute and storage resources.
- Work with AWS services such as EC2, S3, EBS, FSx, VPC, Cloud Watch, and related technologies.
- Support integration between on-premises HPC/storage infrastructure and cloud resources.
- Assist with cloud architecture, resource provisioning, performance optimization, and cost management.
- Support containerized workloads using Docker.
- Support or assist with Kubernetes platforms such as Amazon EKS, including integration with persistent storage technologies such as Ceph.
- Automate infrastructure provisioning, configuration, monitoring, and maintenance using Bash, Python, or similar technologies.
- Support Git-based infrastructure and CI/CD workflows where appropriate.
- Collaborate closely with bioinformatics scientists, software engineers, laboratory personnel, and technical leadership.
- Translate scientific computing and data requirements into practical infrastructure solutions.
- Perform root-cause analysis for significant infrastructure and performance problems.
- Develop and maintain technical documentation, procedures, and system configurations.
- Evaluate emerging HPC, storage, cloud, and infrastructure technologies.
- Bachelor's degree in Computer Science, Information Systems, Engineering, or related technical field, or equivalent professional experience.
- 7+ years of experience administering enterprise Linux environments, with strong systems administration and troubleshooting expertise.
- Experience…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).