×
Register Here to Apply for Jobs or Post Jobs. X

Network Operations Center Technician II

Job in Austin, Travis County, Texas, 78716, USA
Listing for: Cirrascale Corporation
Full Time position
Listed on 2026-07-29
Job specializations:
  • IT/Tech
    IT Support, Technical Support, Systems Administrator, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 55000 - 90000 USD Yearly USD 55000.00 90000.00 YEAR
Job Description & How to Apply Below

Cirrascale Cloud Services provides high-performance cloud infrastructure purpose-built for deep learning, generative AI, and large-scale AI inference workloads. We specialize in dedicated GPU cloud solutions tailored to the unique needs of startups, research labs, and enterprise AI teams. Our mission is to accelerate AI innovation by combining powerful hardware with white-glove service and flexible, custom-built environments.

Position Overview

As a Network Operations Technician II at Cirrascale, you will be an integral part of our Operations team, responsible for maintaining the integrity and functionality of our data centers. The ideal candidate will have a strong background working in a 365-days 24/7 Network Operations Center environment. In this role, the NOC Technician will be responsible for monitoring the network via an alarm reporting tool.

NOC Technician will create an outage ticket in Jira (Atlassian). Customers will be notified of the outage and that we are working on a trouble ticket to resolve it. You are the face of our company to our customers, ensuring that every interaction is a best-in-class customer service experience, so being able to demonstrate quality written and oral communication with customers and other stakeholders is of paramount importance.

The primary responsibilities associated with the position include providing technical support to customers who are experiencing problems with their network and mentoring less experienced colleagues. The ideal candidate will have a strong background in SOC 2 compliance, Linux operating systems, a thorough understanding of networking principles, and experience interfacing GPUs remotely and Internet outages.

Key Responsibilities

  • First-line response to alerts and incidents to systems and job failures
  • Demonstrate understanding of GPU nodes and how they are deployed, networked, and clustered within a datacenter
  • Assist customers with ticket triage and basic troubleshooting using the Jira (Atlassian) ticketing system
  • Perform high-level monitoring and troubleshooting on all nodes and network equipmen t
  • Remotely Troubleshoot Installed Servers & GPUs at various global datacenter locations
  • Resolve complex and critical incidents within our datacenter
  • Lead major incident response and coordinatio n
  • Review and optimize existing NOC procedures
  • Work on capacity planning and performance monitorin g
  • Knowledge and experience working with Dell, Super Micro & Lenovo type Servers is highly recommended
  • Perform deep troubleshooting of GPU node failures, job preemption conflicts, and cluster imbalance
  • Analyze alerts for GPU utilization inefficiencies, failed Machine Learning pipelines, or I/O bottlenecks
  • Provide on-site and remote support to resolve urgent technical issues
  • Document system configurations, updates, and inventory, maintaining accurate records of data center assets.
  • Stay current with industry trends, emerging technologies, and best practices in HPC network operations center trends
  • Analyze alerts for GPU utilization inefficiencies, failed Machine Learning pipelines, or I/O bottlenecks
  • Provide on-site and remote support to resolve urgent technical issues
  • Document system configurations, updates, and inventory, maintaining accurate records of data center assets.
  • Stay current with industry trends, emerging technologies, and best practices in HPC network operations center trends

Qualifications

  • 2-4 years of experience in HPC, AI infrastructure, cloud systems, or related
  • Understanding of scripting (Python, Bash, etc.), GPU resource monitoring preferred
  • Solid understanding of HPC datacenter networking principles and experience with network troubleshooting
  • Strong analytical and problem-solving skills, with the ability to work independently and manage multiple tasks
  • Excellent communication skills and the ability to collaborate effectively with the customer and the team. Customer Service is a must
  • Certifications:

    Advanced Linux or any other AI/ML certifications are a huge plus
  • Experience in Datacenter Network Operations
  • Experience with RMAs, logistics, shipping, and receiving a plus
  • It is a plus with experience working in Jira (Atlassian) ticketing system.
  • Proficient in Microsoft 365(Outlook, Word, Excel)
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary