Operations and Support Lead
Listed on 2026-07-27
-
IT/Tech
SRE/Site Reliability, IT Infrastructure, Network Engineer, Systems Engineer
Hydra Host operates mission-critical AI infrastructure where customer success depends on operational excellence. This role combines Customer Support Leadership with Infrastructure Operations Management, serving as the operational hub between customers, engineering, deployment teams, hardware vendors, and AI Factory partners. You'll own customer-facing operational support while building the internal processes that keep our Neo Cloud platform running efficiently. Whether responding to a GPU outage, coordinating infrastructure deployments, managing vendor escalations, improving SLAs, or building scalable operational workflows, you'll ensure both our customers and our infrastructure perform at the highest level.
This is not a traditional support management position. This is an operations leadership role responsible for the daily execution, reliability, and continuous improvement of Hydra Host's AI Factory platform.
Develop and lead Hydra Host's customer support organization supporting:
- Design and build Hydra Host's customer support organization from the ground up
- Enterprise AI customers
- Interface with the Machine Learning engineering teams
- GPU infrastructure customers
- AI Factory operators
- Data center partners
Build a high-performing support organization that delivers exceptional customer experiences while maintaining enterprise-grade service levels.
Own Operational ExcellenceDrive the day-to-day operational health of Hydra Host's Neo Cloud platform by coordinating activities across engineering, infrastructure, vendors, and customer-facing teams. Ensure infrastructure operates reliably while continuously improving operational efficiency.
Lead Major Incident ManagementServe as the Incident Commander during production-impacting events. Coordinate engineering, networking, infrastructure, vendors, and customers to resolve:
- GPU cluster failures
- Network outages
- Hardware failures
- Firmware issues
- Storage performance degradation
- Infrastructure capacity constraints
- Customer-impacting production incidents
Own customer communications throughout incident response while driving rapid resolution and post-incident improvements.
Build Scalable Support OperationsDesign and implement:
- Ticketing systems
- Escalation procedures
- Knowledge management
- Support automation
- AI-powered support tools
- On-call rotations
- Operational playbooks
- Customer communication standards
Partner with deployment, engineering, and data center teams to coordinate:
- Infrastructure deployments
- Rack turn-up
- GPU cluster readiness
- Network activation
- Capacity planning
- Maintenance scheduling
- Production acceptance
- Operational readiness reviews
Own operational relationships with:
- Hardware manufacturers
- GPU vendors
- Data center operators
- Construction partners
- Logistics providers
- Network carriers
- Service providers
Establish, monitor, and improve operational KPIs including:
- SLA compliance
- MTTR
- Incident response times
- Customer satisfaction
- Infrastructure uptime
- Vendor performance
- Capacity utilization
- Operational readiness
Provide executive reporting on operational performance and identify opportunities for continuous improvement.
Cross-Functional LeadershipCollaborate daily with:
- Infrastructure Engineering
- Network Engineering
- Platform Engineering
- Customer Success
- Deployment Program Managers
- AI Factory Partners
- Executive Leadership
- Enterprise Customers
Act as the operational bridge between technical teams and customer-facing organizations.
Build the OrganizationAs Hydra Host grows, you'll help recruit, mentor, and develop the Operations & Support organization while establishing the culture, processes, and operational standards that define world-class infrastructure operations.
Required Qualifications- 5+ years leading customer support, technical operations, infrastructure operations, or service delivery organizations
- Experience supporting enterprise infrastructure, cloud platforms, AI infrastructure, Neo Cloud providers, or large-scale data center environments
- Experience managing production incidents in mission-critical environments
- Strong understanding of servers, networking, storage, and enterprise…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).