Cloud Systems Engineer
Listed on 2026-09-28
-
IT/Tech
IT Infrastructure, Unix/Linux, Systems Administrator
We are seeking a Cloud Systems Engineer to support and operate large-scale AI and high-performance computing (HPC) environments. This role will be responsible for the deployment, maintenance, performance, and lifecycle management of GPU-accelerated compute infrastructure that powers critical AI, machine learning, and data-intensive workloads.
The ideal candidate is a hands-on infrastructure professional with strong Linux administration skills, deep hardware troubleshooting experience, and expertise supporting enterprise-class compute platforms. This individual will work closely with infrastructure, networking, storage, and AI engineering teams to ensure the reliability, scalability, and operational excellence of our AI infrastructure.
Responsibilities:
AI Infrastructure Operations
- Deploy, configure, and maintain GPU-accelerated compute infrastructure.
- Manage operating system, firmware, BIOS, BMC, driver, and software lifecycle updates.
- Monitor system health, performance, utilization, and capacity across AI infrastructure environments.
- Support infrastructure utilized for AI model training, inference, and data processing workloads.
- Develop and maintain operational standards, runbooks, and maintenance procedures.
- Participate in on-call support and incident response activities.
- Administer enterprise Linux environments, including Ubuntu and Red Hat-based distributions.
- Perform system patching, hardening, and operating system lifecycle management.
- Troubleshoot operating system, kernel, storage, networking, and application-level issues.
- Develop automation to streamline deployment, monitoring, and operational processes.
- Support security and compliance initiatives across AI infrastructure platforms.
Hardware and Datacenter Operations
- Install, configure, maintain, and troubleshoot enterprise compute hardware.
- Diagnose and resolve issues involving GPUs, CPUs, memory, storage, power, and networking components.
- Perform firmware upgrades and hardware lifecycle management activities.
- Coordinate hardware replacements, vendor support engagements, and warranty services.
- Participate in rack-and-stack deployments, datacenter expansions, and technology refresh projects.
- Maintain accurate asset inventories and operational documentation.
- Support high-performance networking technologies, including Ethernet and Infini Band environments.
- Collaborate with networking, storage, cloud, and AI engineering teams on infrastructure design and operations.
- Assist with scalability, resiliency, and performance optimization initiatives.
- Perform root-cause analysis of infrastructure failures and develop preventative measures.
- Other duties as assigned.
Required Qualifications
Experience
- 3-5 years of Linux systems administration experience in production environments.
- 3-5 years of experience supporting enterprise server infrastructure.
- Experience supporting large-scale compute environments, HPC platforms, AI infrastructure, or GPU-enabled systems.
- Experience performing hardware diagnostics, firmware management, and lifecycle maintenance.
- Experience working within datacenter operations environments.
- Bash, Python, Power Shell, or similar scripting languages
- Operating system performance tuning and monitoring
- Storage and networking fundamentals
- Experience with infrastructure monitoring and observability platforms, ex Grafana.
- Hardware and firmware lifecycle management
Please note that sponsorship of new applicants for employment authorization, or any other immigration-related support, is not available for this position at this time.
- Collaborate with outstanding people: We hire only the best. Our standards are high and our employees enjoy working alongside other high achievers.
- Make an immediate impact: New employees can expect to be given real responsibility for bringing new technologies to the marketplace. You are empowered to perform as soon as you join the team!
- Gain well rounded experience: offers a diverse and dynamic environment where you will get the chance to work directly with executives and develop expertise across multiple areas of the business.
- Community and Camaraderie: One of our core values is to 'Keep It Fun,' which to us means fostering a strong sense of community. Our culture is built on collaboration and connection, where we celebrate our successes and believe that a positive, engaging environment is key to doing our best work.
- values working together and collaborating in person. Our employees work from the office 4 days a week.
is the leading platform for intelligently connected…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).