Storage Operations Engineer
San Diego, San Diego County, California, 92189, USA
Listed on 2026-08-18
-
IT/Tech
Cloud Computing: Infrastructure & Operations, IT Infrastructure, Systems Engineer, SRE/Site Reliability
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Storage Operations Engineer based in the United States.
This is a remote engineering role focused on building, operating, and continuously improving large-scale cloud storage infrastructure.
You’ll work extensively with software-defined storage, particularly Ceph-based environments, to improve stability, scalability, performance, and cost efficiency.
The role combines hands-on systems engineering with automation, monitoring, architecture, and operational excellence.
You’ll collaborate closely with networking, operations, and engineering teams to maintain highly available storage services and expand capacity.
You’ll also evaluate emerging technologies and develop new storage capabilities that can improve the efficiency of cloud infrastructure.
This is an opportunity to take meaningful ownership in a fast-growing, technically demanding environment while shaping reliable infrastructure at scale.
- Operate, maintain, and continuously enhance distributed cloud storage environments, with a strong focus on stability, scalability, functionality, performance, and cost efficiency.
- Contribute to the architecture and design of Ceph-based storage platforms and help make sound technical decisions around evolving infrastructure.
- Develop automation frameworks, scripts, and operational improvements that reduce manual work and increase reliability.
- Optimize monitoring, metrics collection, alerting, and observability systems to proactively identify and resolve infrastructure issues.
- Monitor storage infrastructure and help maintain high availability, reliability, and consistent service performance.
- Partner with networking and operations teams on capacity expansion, maintenance windows, and infrastructure lifecycle activities.
- Research emerging storage technologies, operational methodologies, and advanced Ceph capabilities to improve infrastructure efficiency.
- Contribute to technical documentation and internal knowledge-sharing resources.
- Work effectively across engineering and operational teams, communicating technical information clearly and contributing to meetings and knowledge-transfer activities.
- Manage assigned work independently while collaborating effectively with project stakeholders and team members.
- 3+ years of hands-on experience configuring, deploying, and managing distributed systems, with strong understanding of distributed systems concepts and foundational data structures.
- Strong Linux administration experience across environments such as CentOS, RHEL, Debian, or Ubuntu.
- Practical knowledge of cloud hosting, object storage, and block storage concepts.
- Experience with performance tuning in virtualized environments and familiarity with server-class hardware, including IPMI, SATA, and NVMe storage.
- Strong scripting and automation capabilities using Python, Bash, PHP, or similar languages.
- Experience with infrastructure automation tools such as Puppet, Ansible, Chef, or Salt.
- Experience with Ceph or other software-defined storage technologies is highly valuable.
- Familiarity with logging, monitoring, metrics, and time-series technologies such as Prometheus, collectd, ELK, Grafana, Graylog, or Graphite is a plus.
- Strong problem-solving and troubleshooting abilities, with the capacity to diagnose complex infrastructure issues systematically.
- Excellent communication and collaboration skills, with the ability to work effectively across technical teams.
- Strong time-management and organizational skills, combined with the ability to prioritize and operate independently.
- A proactive mindset and willingness to research new technologies, improve existing systems, and take ownership of operational challenges.
- Salary: $80,000–$100,000 annually, with final compensation determined by experience, skills, location, and applicable laws.
- Remote work: Fully remote position for eligible employees based in the United States.
- Healthcare: 100% company-paid medical, dental, and vision insurance premiums for employees.
- Retirement: 401(k) plan with a 100% company match up to 4%, with…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).