SME Data Center-Operation
A Subject Matter Expert (SME) in Data Center Operations acts as the ultimate technical authority for a facility's critical infrastructure. They do not just monitor equipment—they own the complex engineering strategies, standard operating procedures, and emergency response plans that keep a data center running with zero downtime.
While a Commissioning Manager focuses on building and testing new systems, the Operations SME focuses on maintaining, optimizing, and troubleshooting live, active systems.
Job SummaryThe Data Center Operations SME provides high-level technical expertise to sustain maximum uptime for critical electrical, mechanical, and control systems. They serve as the final tier of technical escalation for complex infrastructure failures, drive root-cause analysis (RCA), lead capacity planning, and ensure operations comply with global safety, efficiency, and sustainability standards.
- High-Level Troubleshooting:
Serve as the primary technical point of escalation for complex, unresolved incidents affecting critical systems (e.g., erratic UPS behavior, chiller plant short-cycling). - MOP & SOP Ownership:
Write, review, and approve high-risk Method of Procedures (MOPs), Standard Operating Procedures (SOPs), and Emergency Operating Procedures (EOPs) to ensure zero human error during maintenance. - Root Cause Analysis (RCA):
Lead forensic engineering investigations following critical infrastructure anomalies or near-misses, producing detailed technical reports and preventive action plans. - Capacity & Efficiency Management:
Analyze facility power usage effectiveness (PUE), optimize airflow management, and track electrical/thermal capacity limits as IT loads fluctuate. - Vendor & Lifecycle Management:
Audit external vendor service contracts, oversee major lifecycle overhauls (e.g., generator rebuilds, UPS battery replacements), and ensure spare parts inventory meets strict SLA requirements. - Training & Mentorship:
Design and deliver technical training programs to upskill site technicians, ensuring the operational team is certified to handle emergency failovers.
- Education:
A Bachelor’s degree in Electrical or Mechanical Engineering, or intensive military training in critical environments (e.g., Navy Nuclear program, USACE). - Experience:
Minimum 7+ years of hands-on operations experience inside live, high-availability critical environments (Data Centers, Semiconductor Fabs, or Hospitals). - Technical Mastery:
Deep, localized expertise in either Critical Power (e.g., 11kV/415V distribution, dual-source static switches, rotary/static UPS) OR Critical Cooling (e.g., chilled water loops, CRAH units, hot/cold aisle containment). - Systems Fluency:
Advanced proficiency in navigating and modifying Building Management Systems (BMS), Electrical Power Monitoring Systems (EPMS), and DCIM platforms.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).