NOC Engineer
Job in
Miami, Miami-Dade County, Florida, 33222, USA
Listed on 2026-08-22
Listing for:
Axe Compute
Full Time
position Listed on 2026-08-22
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Job Description & How to Apply Below
Axe Compute is seeking NOC Engineers who do more than watch a screen. This role owns customer experience for live GPU cluster deployments, using automation, AI tools, and continuous process improvement to drive SLA performance, not just monitor it. We're looking for people who build tools, write and improve their own runbooks and SOPs, and treat every shift as an opportunity to make the NOC better than they found it.
ROLEAT A GLANCE
- Mandate:
Monitor live GPU clusters, drive incident response, and continuously improve the tools, processes, and SOPs the NOC runs on. - Scope: 24/7 shift-based monitoring and incident response, tool-building and automation, runbook/SOP creation and maintenance, third-party coordination, SLA metric improvement.
- Key Outcomes: SLA compliance, a measurably improving NOC (fewer repeat incidents, faster resolution, better tooling), strong customer experience outcomes.
- Monitor cluster health, power, cooling, and network status across live deployments, triaging and resolving incidents within SLA.
- Escalate to Tier 3 (OEM or data center operator) only when an issue genuinely requires vendor-level support.
- Build and improve internal tools and automation to reduce manual, repetitive work, using AI-enabled workflows wherever they make the team faster or more accurate.
- Identify opportunities to automate recurring monitoring, triage, or reporting tasks rather than performing them manually shift after shift.
- Create and maintain runbooks and standard operating procedures, updating them based on real incidents rather than leaving them static.
- Contribute to after-action reviews and turn recurring issues into permanent process or tooling fixes.
- Own the customer experience of every incident, communicating clearly and proactively rather than treating tickets as a checklist.
- Work directly with third parties (data center operators, OEMs, network providers) as needed to resolve client-impacting issues.
- Track and actively work to improve SLA metrics, using data and software tooling to identify where performance is slipping before it becomes a client-facing problem.
- 1+ years in a NOC, network operations, or infrastructure monitoring role. This role can be junior, but not passive.
- Demonstrated interest or experience in automation, scripting, or tool-building, not just following existing procedures.
- Comfort working rotating shifts, including nights/weekends, as part of a 24/7 coverage model.
- Familiarity with monitoring tools (e.g., Datadog, Grafana, Pager Duty) and a genuine curiosity about AI-enabled operations tools.
- Experience monitoring GPU or high-performance computing infrastructure.
- Scripting ability (Python, Bash, or similar) to build or modify automation.
- Additional languages beyond English are a plus for coordinating with clients and global vendors.
- Multiple shift options available to accommodate different timezones and flexible working hours.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×